Astrostatistics News
Issue 9, Oct 2026
Issue Editors: Jessi Cisewski-Kehe, David W. Hogg, Vinay L. Kashyap, Aneta Siemiginowska
Astrostatistics News (AN) is a newsletter designed to inform, promote, cultivate, and inspire the astrostatistics community.
2027 Astrostatistics Student Paper Competition
Special Issue on Statistics for Astronomical Imaging Data
ASA/AIG 2026 Student Paper Competition Results
From Cosmos to Dempster, Bypassing Bayes?
Where to get Astronomical Data: Sunspot Numbers
Jargon Dictionary: Calibration, because, almost surely
Astrostatistics Events: STAMPS, IAU-IAA Seminars
Other News: Job Opportunities, Content Suggestions, Website, Subscription
David van Dyk (Imperial College London)
The Astrostatistics Interest Group of the American Statistical Association is pleased to announce its ninth annual Astrostatistics Student Paper Competition. From all the papers submitted to competition, up to five finalists will be chosen to select their work in a special astrostatistics session at the 2027 Joint Statistical Meetings in Chicago, Illinois. Each of the five finalists will receive a $100 cash prize, and the winner (announced at the SPC session) will receive an additional $400.
Papers will be judged on the clarity and framing of the scientific problem, the novelty of the proposed statistical methods, thoroughness of the results and their comparisons with prior approaches, and the overall presentation and cross-disciplinary impact of the work.
The application deadline for the 2026 Astrostatistics Student Paper Competition is Tuesday, December 1, 2026, at 11:59 pm PST.
Application details and eligibility requirements can be found here.
Please share this announcement with any students writing papers related to the development and/or application of statistical methodologies in astronomy, astrophysics, or cosmology. As always, the AIG strongly encourages applications from students from groups underrepresented in statistics and/or astronomy!
Aneta Siemiginowska
Website: https://think.taylorandfrancis.com/special_issues/astronomical-imaging-data/
The open access Journal of Statistics and Data Science in Imaging (SDSI), published on behalf of the American Statistical Association, aims to address a significant gap in current outlets for researchers working on statistical methods and data science techniques for imaging data analysis. The primary aim is to serve as a forum for discussing methodological challenges encountered in the analysis of imaging data and for presenting statistically sound solutions. The issue editors are David Stenning, Yang Chen, Vinay Kashyap, Thomas Loredo, Aneta Siemiginowska, and Marina Vannucci.
The special issue on Statistics for Astronomical Imaging Data is accepting new manuscripts before December 31, 2026. All the contributions are peer-reviewed. This special issue aims to showcase recent advances in innovative astrostatistical methodologies and computation which have been developed for processing and analyzing imaging data in astronomy, and looks ahead into the future of astrostatistics.
There are no publication charges for papers accepted for publication in this special issue.
Note: Manuscripts due by December 31, 2026
More details are available on the site.
Astrostatistics innovations of the present are highlighted in this section.
An annual competition to identify innovative student-led papers is run by the American Statistical Association’s Astrostatistics Interest Group; details on the 2027 competition can be found above under “2027 Astrostatistics Student Paper Competition.” The finalists present their work in a special session at the Joint Statistical Meeting. A summary of the papers of the 2026 finalists are provided below, written by the finalists.
Winner: Konstantin “Kosio” Karchev (Imperial College/University of Barcelona)
Karchev, Konstantin, Roberto Trotta, and Raúl Jiménez. "CIGaRS I: combined simulation-based inference from type Ia supernovae and host photometry." Nature Astronomy (2026): 1-16.
Type Ia supernovae (SNae Ia) are essential tools for measuring the expansion of the Universe. However, their cosmographic use remains largely empirical—through a process known as standardisation, which relates observable properties like duration (stretch) and colour to absolute brightness and thence, distance—and their progenitor mechanism largely unresolved. Investigating them in the context of their host galaxies thus has the potential to both improve the precision of SN Ia distance estimates and shed light on the physical processes that govern them, but this requires: disentangling intrinsic colour and brightness scatter from extrinsic (e.g., dust-related) attenuation, measurement noise, progenitor-related effects, cosmological distance, and selection biases, all while inferring the host’s stellar composition and chemical properties (metallicity). This is all but impossible to achieve with traditional Bayesian techniques (MCMC) but a breeze for simulation-based inference, provided a suitable forward model and an appropriately structured neural network.
This work presents the first unified Bayesian hierarchical forward model for galaxy evolution (stellar population synthesis, mass, dust, and metallicity through cosmic time) and type Ia supernovae, connected on two levels: first, through an occurrence model (a delay-time distribution, DTD) deriving the presence of a SN Ia to the stellar population of its host; and second, via a flexible parameterized dependence of intrinsic SN Ia brightnesses on the metallicities and ages of their progenitors, allowing also for an ad-hoc “mass step.” We use it to train a bespoke conditioned deep set neural ratio estimator to simultaneously derive marginal posteriors on parameters of the DTD, standardisation coefficients (including metallicity/age/mass-related), cosmology, and all object-specific latent variables (e.g., stellar mass and redshift) only from broadband host photometry and SN Ia summary statistics. We test the pipeline on mock data from the Rubin Observatory’s LSST with realistic selection effects and demonstrate competitive cosmological performance and state-of-the-art photometric redshift estimation. CIGaRS paves the way to full cosmological utilisation of LSST’s hard-earned data and to studying SNae Ia as astrophysical objects in their own right and in the context of their host galaxies, without methodological compromises.
Finalist: Michael Poon (University of Toronto)
Poon, Michael, Marta L. Bryan, Hanno Rein, Jiayin Dong, Joshua S. Speagle, and Dang Pham. "Early evidence for isotropic planetary obliquities in young super-Jupiter systems." The Astrophysical Journal Letters 994, no. 2 (2025): L48.
This decade has seen the first measurements of extrasolar planetary obliquities, characterizing how an exoplanet's spin axis is oriented relative to its orbital axis. These measurements are enabled by combining projected rotational velocities, planetary rotation periods, and astrometric orbits for directly-imaged super-Jupiters. To test whether these super-Jupiters form more like scaled-up planets or scaled-down stars, we develop a Bayesian hierarchical framework to infer their population-level obliquity distribution. Using a single-parameter Fisher distribution, we compare two models: a planet-like formation scenario (κ = 5) predicting moderate alignment, versus a brown dwarf-like formation scenario (κ = 0) predicting isotropic obliquities. Based on a sample of four young super-Jupiter systems, we find early evidence favoring the isotropic case with a Bayes factor of 15, consistent with turbulent fragmentation.
Finalist: Joseph Salzer (University of Wisconsin-Madison)
Salzer, Joseph, Jessi Cisewski-Kehe, Eric B. Ford, and Lily L. Zhao. "Searching for Low-mass Exoplanets amid Stellar Variability with a Fixed Effects Linear Model of Line-by-line Shape Changes." The Astronomical Journal 170, no. 3 (2025): 179.
A star's spectrum is a continuum characterized by hundreds of narrow absorption lines, each centered at a wavelength set by a particular atomic transition. A planet's gravitational pull moves the star toward and away from Earth periodically, Doppler-shifting every line by the same amount without altering any line's shape. Detecting that wobble allows astronomers to infer the presence of a planet and characterize its mass via the Radial Velocity (RV) Method. For an Earth analog around a Sun-like star the shift is roughly 0.1 m/s. However, stellar variability produces apparent velocity shifts while deforming the line profiles. Line shape is, therefore, a proxy for contamination that is blind to the planet. We obtain daily-averaged solar spectra from the NEID spectrograph, which includes 778 lines over 330 days. All known planetary contributions are removed, so the true radial velocity is exactly zero m/s and the error against ground truth is directly computable.
The model is a two-way fixed effects regression, applied to a panel whose units are spectral lines and whose time periods are observation nights. Each line-by-line RV is modeled as the sum of a line fixed effect, a time fixed effect, and a set of line shape covariates carrying slopes that are estimated separately for every line. The time fixed effects are the quantity of interest, and the cleaned RV time series is simply a linear combination of those estimated terms. Conditioning on line shape variations in the proposed model reduces the RV RMS errors from 1.722 to 0.403 m/s (approximately 76% reduction).
Finalist: Ana Sofia Uzsoy (Harvard University)
Uzsoy, Ana Sofía M., Andrew K. Saydjari, Arjun Dey, Anand Raichoor, Douglas P. Finkbeiner, Eric Gawiser, Kyoung-Soo Lee et al. "Bayesian Component Separation for DESI LAE Automated Spectroscopic Redshifts and Photometric Targeting." The Astrophysical Journal 999, no. 2 (2026): 241.
Lyman-Alpha Emitters (LAEs) are young, star-forming galaxies with significant emission from Lyman-Alpha, an emission line of neutral hydrogen. The Dark Energy Spectroscopic Instrument (DESI) has observed thousands of LAEs; however, the sharp Lyman-Alpha line can often be confused with residual sky emission lines and the DESI data reduction pipeline does not have an automated method for determining LAE redshifts, leaving them to be estimated only with visual inspection (VI). We present a method leveraging Bayesian component separation to separate the Lyman-Alpha line from the sky and noise components and automatically determine redshifts in ~0.4 CPU seconds per galaxy. We use real spectra to create data-driven covariance matrices to serve as priors for each component, and assign each galaxy the redshift that minimizes the change in chi-squared, achieving > 90% agreement with VI redshifts. The value and curvature of the chi-squared minimum indicate redshift confidence and uncertainty, and we ensure these are well-calibrated using injection-recovery tests.
Finalist: Matthew O’Callagan (University of Cambridge)
O’Callaghan, Matthew, Kaisey S. Mandel, and Gerry Gilmore. "Data-driven dust inference at mid-to-high Galactic latitudes using probabilistic machine learning." Monthly Notices of the Royal Astronomical Society 546, no. 2 (2026).
Mapping the diffuse interstellar medium (ISM) and its small-scale variations at mid-to-high Galactic latitudes is essential for a broad range of astronomical and cosmological applications, such as correcting for dust foregrounds in cosmic microwave background (CMB) experiments. However, charting these fine-scale structures in a bias-free manner remains a significant challenge. Traditional reddening-based dust maps typically rely on deterministic functions or theoretical ab initio stellar evolution models to define a baseline zero-extinction sample of stars. These methods are highly susceptible to modelling errors, strong prior assumptions, and stellar parameter degeneracies, which can introduce artificial structural biases into the resulting dust maps and falsely identify genuinely zero-extinction stars as extinguished. This is all but impossible to properly constrain with traditional functional fits, but highly suited for probabilistic generative modelling.
In our paper, we present FLOWER (FLOW-based Estimation of Reddening), a novel, data-driven method that models the photometric colour-magnitude relations of zero-extinction stars as a probability distribution. Using conditional normalising flows, the model learns these intrinsic relations conditioned solely on a star's Galactic cylindrical coordinates (absolute perpendicular height and radial distance). By learning this distribution probabilistically rather than deterministically, the model naturally accounts for observational scatter, astrophysical variations like binary star systems, and selection effects (such as the Malmquist bias) without relying on complex stellar labels. The authors use this forward model to simultaneously derive Bayesian marginal posteriors on line-of-sight photometric dust extinction. Tested on real data from Gaia, Pan-STARRS, and 2MASS, the pipeline successfully recovers unbiased posteriors on synthetically extinguished stars and accurately traces known single- and multi-cloud line-of-sight dust structures in mid-latitude calibration regions that previous mapping methods failed to resolve. Ultimately, FLOWER paves the way for probabilistically robust, high-resolution 3D mapping of the dusty ISM without the systemic biases and overly confident prior assumptions of traditional methodologies.
Xiao-Li Meng (Harvard)
Even if there are infinitely many universes, we can access only one. Taken seriously, every inference about our cosmos rests on a sample of size one, a statistician’s nightmare. Apparent replications come from coarse graining: in full detail, nothing looks or behaves like anything else, deterministically or statistically. Thinking statistically does not escape this; we manufacture "replications"—and hence more samples—by ignoring high-resolution signals we deem irrelevant. Whatever region we study, large or small, is our cosmos, and it always contains something we do not understand, which is why we study it. Bayes is a standard way forward whenever our data or information is too coarse, whatever the sample size might be (or mean). But there is no free lunch. Bayesian calculations demand fully specified probabilities, which inevitably include procedural assumptions, made up not because we believe them but because the procedure requires them. When my astrophysicist collaborators only tell me that a power-law parameter α lies between 0.5 and 1.5, I must supply the rest of the prior for α to apply Bayes theorem. The reflexive uniform prior is far from total ignorance: it asserts that every permissible value is equally likely—that is a very strong prior! Worse, a uniform α does not imply a uniform α³, yet why should a one-to-one transformation of α matter under total ignorance? Can we use only what we know, letting low-resolution information, whether data, priors, or models, yield low-resolution answers? That is the aim of imprecise probability, a concept more familiar than it sounds. Invited to a weekend barbecue, you might put your chance of attending at "60 to 80 percent," expressing uncertainty about your own uncertainty: a lower and an upper probability. The trouble starts with updating. To obtain the probability of A given that B has occurred, Bayes rule divides the probability of "A and B" by that of B, but with imprecise probability, both numerator and denominator come with two numbers. Dividing upper by upper and lower by lower violates a basic logical requirement: the weakest case for an event must equal one minus the strongest case against it, the so-called conjugacy identity. Three main remedies exist: Dempster's rule, which uses only upper probabilities plus the conjugacy identity; the Geometric rule, which uses only lower probabilities plus the conjugacy identity; and generalized Bayes, which maximizes and minimizes the ratio over all distributions consistent with the low-resolution information we have; see Gong and Meng (2021) for an overview and for details of what follows.
Each has appeals and anomalies, some fatal in my opinion. I am very nervous about applying the Geometric rule or generalized Bayes, because neither can ever escape total ignorance, however strong the data. Why would anyone want an updating rule that refuses to learn, especially when one is completely ignorant? That naturally leaves Dempster's rule, a generalization of Bayes rule (but not the generalized Bayes rule!) to imprecise probability. It is named after my Harvard colleague Arthur Dempster (1929–2026), best known for the Dempster–Shafer theory of belief functions and for the EM algorithm, who sadly passed away early this year. His rule does learn, but as an optimist (and hence escapes the trap of total ignorance). In the classic three-prisoners puzzle, merely contemplating a question to the guard raises a prisoner's survival chance from 1/3 to 1/2, whatever the answer: a sure gain. The pessimistic Geometric rule concludes zero chance of survival: a sure loss. Generalized Bayes is the most conservative, declaring anything between 0 and 1/2 permissible. This is the problem of dilation: merely contemplating how the guard might respond dilates the 1/3 chance of survival to anywhere from 0 to 1/2. So can we bypass Bayes when facing low-resolution information? Yes, but not the underlying dilemma. Assumptions about priors, models, or data are often unverifiable, and if we adopt procedural assumptions to plough through, we had better be prepared for their consequences. Imprecise probability avoids such assumptions, but we must then live with the anomalies of whichever rule we choose, and declare ourselves optimist, pessimist, or conservatist. Which is the lesser evil for you, and why?
Reference: Gong, R. and Meng, X.-L. (2021). Judicious judgment meets unsettling updating: dilation, sure loss, and Simpson's paradox (with Discussion and Rejoinder). Statistical Science 36(2), 169–214.
Most astronomical data are public. Large datasets are a mere tempting click away, and it is easy for the unwary to trip over details. Here we will highlight some sources of data which are statistician friendly and whose maintainers have made a significant effort to document their contents.
This list of astronomical data sources is maintained at our website,
https://www.astrostatisticsnews.com/data
Sunspot numbers are a favorite of statisticians, because the data are the longest duration continuous astronomical dataset in existence and are a treasure trove of statistical correlations and trends. The physics of the sunspot cycle is complicated, but the statistical analyses are accessible and approachable. Yearly SSNs have been available starting from 1700, from deep within the Maunder Minimum, monthly compilations since 1749, and daily numbers from 1818. As one would expect from a multi-generational enterprise, the character of the data, how it is measured, and even what a sunspot means has changed over the centuries. The estimated SSNs are now curated by the Royal Observatory of Belgium, and the tables themselves are available at https://www.sidc.be/SILSO/datafiles. The numbers were recalibrated c.2014, and this recalibrated dataset is even available in R, post v4.5.0. Monthly sunspot numbers from 1749 to 1983 are available in the R datasets library under the name sunspots, while (recalibrated) data through October 2024 can be found in the dataset called sunspot.month.
–VLK/JCK
General definitions of astronomy or statistical terms are included in this section.
A list of defined terms is maintained at our website,
https://www.astrostatisticsnews.com/dictionary
The International Astronomical Union is constructing a glossary of common astronomy terms (see https://astro4edu.org/resources/glossary/search/). Here we plan to build up a similar dictionary, focusing on both statistics and astronomy jargon.
If you have comments, questions, concerns, edits, or terms you would like included please let us know at astrostatisticsnews@gmail.com.
Calibration is understood differently in Astronomy (and Physics) and in Statistics. In Astronomy, it is part of a forward process that translates physics-based quantities (like energy flux) to the output of actual instruments. A common usage of “calibration” in Statistics, which appears to have borrowed the term from Chemistry, is generally part of an inverse regression process where a response can be mapped onto a less easily obtained derived feature. In principle, both terms imply a pre-determined mapping from an observational to an inferential space, but go about it very differently. A good description of the difference is in Footnote 8 of Tak et al. 2024 (ApJS 275, 30).
–VLK/JCK
Much to my surprise, I keep encountering people who have never seen this symbol. It was common to find it in math proofs, but many seem to be unfamiliar with it generally. Most are aware of others in the same set: ∴ (therefore), ∃ (there exists), ∈ (element of), etc., but somehow ∵ has fallen through the cracks.
–VLK
Recall that random variables are real-valued functions on a sample space, and each random variable assigns a real number to every possible outcome in the sample space. Now consider a sequence of random variables XN, for N=1, 2, … . We say XN converges almost surely to a random variable X, if for every possible outcome in the sample space, except for a set of probability zero, XN converges to X.
This is stronger than convergence in probability that merely asserts that, for any positive ε, the probability that XN and X differ by more than ε converges to zero as N goes to infinity.
–EBH
A list of job opportunities is maintained at our website, astrostatisticsnews.com/job-opportunities.
A list of events is maintained at our website, astrostatisticsnews.com/upcoming-events.
STAMPS Seminar Series
STAtistical Methods for the Physical Sciences Research Center (STAMPS@CMU)
https://www.cmu.edu/dietrich/statistics-datascience/stamps/index.html
launched the seminar series on September 20, 2024.
Talks are open to everyone who registers on the web site:
https://www.cmu.edu/dietrich/statistics-datascience/stamps/events/webinars/index.html
IAU - IAA Astrostatistics and Astroinformatics Seminars
Monthly Virtual online
Details: https://sites.google.com/view/iau-iaaseminar-new/home
Schedule: https://sites.google.com/view/iau-iaaseminar-new/schedule?authuser=0
This international online seminar is an initiative of the International Astrostatistics Association and the IAU Astroinformatics and Astrostatistics Commission. It focuses on statistical analysis and data mining of astronomical data. The seminar is run on Zoom monthly on second Tuesdays alternating between Europe-America and Australasia-Europe time zone instances. The standard seminar times are 8:00 UTC and 16:00 UTC. Please check the exact time and time differences with your timezone.
If you have ideas for AN content, please send a message to astrostatisticsnews@gmail.com.
Ideas may include relevant astrostatistics papers/data/code, visualizations, upcoming events, job postings, format or commentary suggestions, etc.
See astrostatisticsnews.com for more information such as past issues, lists of astrostatistics references and societies.
To subscribe to Astrostatistics News, go to https://groups.google.com/g/astrostatistics-news and select the “Join group” button. You will need to be logged into your Google account to join the group.
Please forward this information to anyone who may be interested!