Previous seminars 2025

Spring

Seminar 2025-01-29: Bin-Conditional Conformal Prediction of Fatalities from Armed Conflict

Speaker David Randahl, Department of Peace and Conflict Research, Uppsala University

Topic Bin-Conditional Conformal Prediction of Fatalities from Armed Conflict

Abstract Armed conflict forecasting is an important area of research that has the potential to save lives and prevent suffering. However, most existing forecasting models provide only point predictions without any individual-level uncertainty estimates. In this paper, we propose a novel extension to conformal prediction which allows users to obtain individual-level prediction intervals for any arbitrary prediction model and which maintains a user-specified level of coverage across bins of values defined by the user. We then apply this bin-conditional conformal prediction model to forecast fatalities from armed conflict. The results show that the bin-conditional conformal prediction model provides well-calibrated uncertainty estimates for the predicted number of fatalities. The bin-conditional method outperforms the standard conformal prediction method with respect to thecalibration of coverage rates across different values of the prediction outcome, but at the cost of wider prediction intervals.

 

Seminar 2025-02-05: Legal and Ethical Implications of Transformer-Assisted Hate Crime Classification and Estimation in Sweden

Speaker Hannes Waldetoft, Department of Statistics at Uppsala University

Topic Legal and Ethical Implications of Transformer-Assisted Hate Crime Classification and Estimation in Sweden

Abstract Hate crimes, driven by biases against specific demographic groups, harm not only individuals but undermine the security, trust, and cohesion of entire communities. Accurately identifying such crimes remains a significant challenge due to under-reporting, limited training, and the complexity of determining bias motivations. In this paper, we analyze the results of a transformer-based classification model developed to improve the precision of hate crime statistics and identification in Sweden. Empirical results indicate the model outperforms traditional manual police classification of hate crimes, achieving higher precision across various crime types and regions. We further disaggregate performance to pinpoint persistent challenges and highlight categories where both human and machine decision-makers struggle. While the model focuses on statistical estimation rather than direct case-level decision-making, we discuss the broader implications of algorithmic transparency, accountability, and explainability. Ultimately, this research illustrates how transformer-based neural networks can responsibly bolster the detection and understanding of hate crimes, informing policies to better protect vulnerable communities

 

Seminar 2025-02-26: Decomposing Global Bank Network Connectedness: What is Common, Idiosyncratic and When?

Speaker Luca Margaritella, Department of Economics, Lund University

Topic Decomposing Global Bank Network Connectedness: What is Common, Idiosyncratic and When?

Abstract We propose a novel approach to estimate high-dimensional global bank network connectedness in both the time and frequency domains. By employing a factor model with sparse VAR idiosyncratic components, we decompose system-wide connectedness (SWC) into two key drivers: (i) common component shocks and (ii) idiosyncratic shocks. We also provide bootstrap confidence bands for all SWC measures. Furthermore, spectral density estimation allows us to disentangle SWC into short-, medium-, and long-term frequency responses to these shocks. We apply our methodology to two datasets of daily stock price volatilities for over 90 global banks, spanning the periods 2003-2013 and 2014-2023. Our empirical analysis reveals that SWC spikes during global crises, primarily driven by common component shocks and their short-term effects. Conversely, in normal times, SWC is largely influenced by idiosyncratic shocks and medium-term dynamics.

 

Seminar 2025-03-05: Are people with chronic pain more diverse than we think? An investigation of ergodicity

Speaker Felicia Sundström, Department of Psychology, Uppsala University

Topic Are people with chronic pain more diverse than we think? An investigation of ergodicity

Abstract This study investigates whether data from people with endometriosis (n = 58) and fibromyalgia (n = 58) exhibit what is called “ergodicity,” meaning that results from analyses of aggregated group data can be used to support conclusions about the individuals within the groups. The variables studied here are commonly investigated in chronic pain: pain intensity, pain interference, depressive symptoms, psychological flexibility, and pain catastrophizing. Data were collected twice daily for 42 days from each participant and analyzed in two ways: as separate cross-sectional group studies using the timepoints as the separate data sets (between-person) and as individual longitudinal studies using each person's time series data (within person). To confirm ergodicity, the results from the two analyses should agree. However, this is not what was observed in several respects. The between-person data showed substantially less variability compared with within-person data. This was evident in both the summary statistics involving single variables and in the correlational analyses. Overall, between-person correlations were relatively restricted in range, while within-person correlations varied widely. These findings have potentially profound implications for the field of chronic pain research. Because ergodicity was not found, this raises doubts around the assumption that aggregated data collected from groups can accurately represent the range of individual experiences in chronic pain. The results advocate for a shift toward inclusion of more individual person-focused approaches as an addition to group-based approaches. This shift could lead to more personalized and effective treatments by better capturing and then clarifying the heterogeneous nature of chronic pain, including the processes that underlie it.

 

Seminar 2025-03-12: Velocities of moving random surfaces

Speaker Krzysztof Podgórski, Department of Statistics, Lund University

Topic Velocities of moving random surfaces

Abstract For a stationary two-dimensional random field evolving in time, one can derive statistical distributions of appropriately defined velocities utilizing a generalization of the Rice formula. The theory can be applied to practical problems where evolving random fields are considered to be adequate models. Examples include changes of atmospheric pressure, variation of air pollution, or dynamical models of the sea surface elevation. In particular, statistical properties of velocities can be obtained both for the sea surface and for the envelope field based on this surface. Additional extension can be obtained by studying three-dimensional geometry of spatial waves. Their statistical distributions can be presented in explicit integral forms for the deep water seas modeled as Gaussian fields. The proposed approach allows for investigation of the effect that shape and directionality of the sea spectrum have on the joint distributions of the size characteristics.

 

Seminar 2025-03-19: Testable implications of outcome-independent MNAR

Speaker Arvid Sjölander, Department of Medical Epidemiology and Biostatistics, Karolinska Institutet

Topic Testable implications of outcome-independent MNAR

Abstract The standard taxonomy for missing data analysis separates missingness mechanisms into “missing completely at random” (MCAR), “missing at random” (MAR) and “missing not at random” (MNAR). Whereas multiple imputation requires MAR for unbiasedness, it is often argued that the simpler complete-case analysis requires the stronger condition MCAR. In this presentation, we will show that a complete-case analysis can be unbiased under a realistic special case of MNAR, which we label outcome-independent MNAR, and we show that multiple imputation is generally biased under this missingness mechanism. This challenges the common assertion that multiple imputation is always preferable to a complete-case analysis, from a bias perspective. We further show that the assumption of outcome independent MNAR can be tested with data. This stands in contrast to MAR, which is fundamentally untestable.

 

Seminar 2025-03-26: A General Design-Based Framework and Estimator for Randomized Experiments

Speaker Fredrik Sävje, Department of Economics, Uppsala University

Topic A General Design-Based Framework and Estimator for Randomized Experiments

Abstract We describe a new design-based framework for drawing causal inference in randomized experiments. Causal effects in the framework are defined as linear functionals evaluated at potential outcome functions. Knowledge and assumptions about the potential outcome functions are encoded as function spaces. This makes the framework expressive, allowing experimenters to formulate and investigate a wide range of causal questions. We describe a class of estimators for estimands defined using the framework and investigate their properties. The construction of the estimators is based on the Riesz representation theorem. We provide necessary and sufficient conditions for unbiasedness and consistency. Finally, we provide conditions under which the estimators are asymptotically normal, and describe a conservative variance estimator to facilitate the construction of confidence intervals for the estimands

 

Seminar 2025-04-02: Preliminary estimation using the LM-principle

Speker Johan Lyhagen, Department of Statistics at Uppsala University

Topic Preliminary estimation using the LM-principle

Abstract Thematic theories are incomplete meaning that they do not give a full description of how to specify a model. Rather they focus on certain aspects of interest which means that there are parts of the model that needs to be empirically decided. This includes lag-lengths in time series analysis, factor structures and correlations amongst errors in SEM, or control variables in causal inference (sensitivity analysis). In SEM there are modification indices, mainly for the purpose of improving the fit of the model, that estimate the increase of the likelihood when relaxing a restriction. Subsequently, one can also derive an estimated parameter change when relaxing a restriction. In this paper we generalise this to relaxing more than one parameter, focusing on the estimated parameter change in the parameters of interest, and derive this in the GLM setting as well as in the traditional SEM. The paper includes theoretical results, Monte Carlo simulations to investigate the small sample properties and empirical examples to show the usefulness for empirical researchers.

 

Seminar 2025-04-09: Supervised learning for repeated measures data

Speaker Martin Singull, Department of Mathematics, Linköping University

Topic Supervised learning for repeated measures data

Abstract Multivariate repeated measures data, which correspond to multiple measurements that are taken over time on each unit or subject, are common in various applications such as medicine, pharmacy, environmental research, engineering, business, finance, etc. In this presentation we will discuss supervised learning, i.e., model fitting and classification, of repeated measurements following a Growth Curve model, which is also known as a bilinear regression model. In the end of the presentation we will also consider some real data example.

 

STATISTICS Review Seminar 2025-05-14: Viktor Eriksson

Speaker Viktor Eriksson, Department of Statistics, Uppsala University

Topic An Extensive Comparison of Small-Sample Properties of Covariance Estimators for MLEs in the Context of ARMA Models

Abstract A common approach for parameter estimation in autoregressive moving-average (ARMA) models is the maximum likelihood estimator (MLE). The inverse Fisher's information matrix (FIM) is consequently an often used estimator for the variance-covariance matrix (VCM) of the MLE. We investigate the small sample properties of five FIM-based estimators for the VCM in a Monte Carlo simulation study. The Box-Jenkins asymptotic estimator and the FIM performed best in respect to the mean squared error (MSE) and relative bias. The observed FIM performed worst in these measures, however it performed best for providing accurate confidence intervals for the ARMA parameters.

 

Fall

 

STATISTICS Seminars Series 2025-09-03: Max Raner

Speaker Max Raner, Department of Mathematics, Uppsala University

Topic False Confidence, Beyond Additivity, and Possib(ilistical)ly a Fusion of the Frequentist and Bayesian

Abstract The quest for the “holy grail” of statistical theory: a posterior distribution without the need for subjective priors, has generated many proposals over the last century. The recent False Confidence Theorem has shown that additive probability measures inevitably risk false certainty, casting doubt not only on these proposed methods, but on probability itself as the language for epistemic uncertainty. A natural way forward is to move beyond additivity, into the framework of imprecise probability. Among these, possibility theory stands out: it is simple, elegant, and uniquely well aligned with the needs of statistical inference, with direct connections to p-values, confidence sets, and hypothesis tests. Possibility offers a language more primitive than probability, and perhaps a bridge between Bayesian and frequentist paradigms—with new territory inbetween. In this talk, I will trace the historical context of this problem, introduce possibility theory and its relation to classical theory, and discuss recent developments and open directions for research.

 

STATISTICS Seminars Series 2025-09-25: David Kohns

Speaker David Kohns, Department of Computer Science, Aalto University

Topic Joint Quantile Shrinkage: A State-Space Approach Toward Non-Crossing Bayesian Quantile Models

Abstract Crossing of fitted conditional quantiles is a prevalent problem for quantile regression models. We propose a new Bayesian modelling framework that penalises multiple quantile regression functions toward the desired non-crossing space. We achieve this by estimating multiple quantiles jointly with a prior on variation across quantiles, a fused shrinkage prior with quantile adaptivity. The posterior is derived from a decision-theoretic general Bayes perspective, whose form yields a natural state-space interpretation aligned with Time-Varying Parameter (TVP) models. Taken together our approach leads to a Quantile-Varying Parameter (QVP) model, for which we develop efficient sampling algorithms. We demonstrate that our proposed modelling framework provides superior parameter recovery and predictive performance compared to competing Bayesian and frequentist quantile regression estimators in simulated experiments and a real-data application to multivariate quantile estimation in macroeconomics.

 

STATISTICS Review Seminar 2025-09-29: Jakob Torgander

Speaker Jakob Torgander, Department of Statistics, Uppsala University

Opponent David Broman, EECS, KTH Royal Institute of Technology

Topic posteriordb: Testing, Benchmarking and Developing Bayesian Inference Algorithms

Abstract The generality and robustness of inference algorithms is critical to the success of widely used probabilistic programming languages such as Stan, PyMC, Pyro, and Turing.jl. When designing a new general-purpose inference algorithm, whether it involves Monte Carlo sampling or variational approximation, the fundamental problem arises in evaluating its accuracy and efficiency across a range of representative target models. To solve this problem, we propose posteriordb, a database of models and data sets defining target densities along with reference Monte Carlo draws. We further provide a guide to the best practices in using posteriordb for model evaluation and comparison. To provide a wide range of realistic target densities, posteriordb currently comprises 120 representative models and has been instrumental in developing several general inference algorithms.

 

STATISTICS Seminars Series 2025-10-08: Jose M. Peña

Speaker Jose M. Peña, Department of Computer and Information Science, Linköping University

Topic Flow IV: Counterfactual Inference In Nonseparable Outcome Models Using Instrumental Variables

Abstract To reach human level intelligence, learning algorithms need to incorporate causal reasoning. But identifying causality, and particularly counterfactual reasoning, remains an elusive task. In this talk, we show the progress we have made on this task by utilizing instrumental variables (IVs). IVs are a classic tool for mitigating bias from unobserved confounders when estimating causal effects. While IV methods have been extended to nonseparable structural models at the population level, existing approaches to counterfactual prediction typically assume additive noise in the outcome. In this talk, we show that under

standard IV assumptions, along with the assumptions that latent noises in treatment and outcome are strictly monotonic, the treatment–outcome relationship becomes uniquely identifiable from observed data. This enables counterfactual inference even in nonseparable models. We implement our approach by training a normalizing flow to maximize the likelihood of the observed data, demonstrating accurate recovery of the underlying outcome function. We call our method Flow IV.

 

STATISTICS Seminars Series 2025-10-15: Gilbert Mutungi

Speaker Gilbert Mutungi, Makerere University, Uganda

Topic Decoding the role of time series features in LSTM forecasting: Evidence from the M4 Dataset

Abstract Long Short-Term Memory (LSTM) networks are a benchmark deep learning model for time series forecasting, yet the factors driving their performance remain unclear. We extract a comprehensive set of time series features from the M4 dataset and assess their influence on LSTM forecast accuracy. Using correlation analysis, Random Forest feature importance, and regression modeling, we evaluate their effect on forecast error. Results show that series length per forecast horizon, skewness, autocorrelation structure, and stationarity strongly predict LSTM performance, while nonlinearity and persistence have little impact. These findings clarify when LSTMs perform well and offer practical guidance for their application in forecasting tasks.

 

STATISTICS Seminars Series 2025-10-22: Pär Stockhammar

Speaker Pär Stockhammar, Department of Statistics, Stockholm University

Topic Econometrics in Swedish policymaking - Applications and reflections from 15 years of practice

Abstract I will talk about how econometrics shapes policy decisions at major Swedish institutions, including the Ministry of Finance, National Institute of Economic Research and the Riksbank. For instance, I'll discuss how macroeconomic forecasts are built and how VAR models have been used to answer policy-relevant questions. The presentation will be very non-technical and hopefully easily accessible also for non-experts. All are welcome!

 

STATISTICS Seminars Series 2025-10-29: Carl Bonander

Speaker Carl Bonander, School of Public Health & Community Medicine, University of Gothenburg

Topic Reproducibility and Data Irregularities

Abstract Over the past decade, there has been a growing focus on reproducibility and research transparency across the quantitative social sciences. Open data policies and replication requirements have made it easier to verify published findings and, in some cases, uncover problems in the underlying data and analyses. In this seminar, I will talk about how these developments have enabled reproducibility work in economics and related fields, and what such work can reveal. I will also present an ongoing academic misconduct investigation and give examples of how statistical and non-statistical methods can help identify data irregularities indicating manipulation and other issues.

 

STATISTICS Seminars Series 2025-11-12: Elisavet Syriopoulou

Speaker Elisavet Syriopoulou, Department of Medical Epidemiology and Biostatistics, Karolinska Institutet

Topic Evaluating mediator interventions for time-to-event outcomes: A causal framework for cancer disparities

Abstract Understanding cancer survival disparities often requires evaluating the impact of potential interventions on mediators that lie between an exposure and an outcome. Causal mediation analysis can be applied in such settings, and it allows exploring such interventions in a systematic way. While traditional approaches have focused on uncovering mechanistic pathways, despite the presence of ill-defined interventions, recent developments have introduced interventional effects that map target trials and focus on evaluating shifts in mediator distributions. I will present some ongoing work in which we extend these interventional effects to settings with time-to-event outcomes and incorporate a relative survival framework. Relative survival is a commonly used measure in cancer epidemiology used to estimate net survival, i.e. the disease-specific survival, without needing to utilise the cause of death information obtained from cancer registers that may be inaccurate or not available. An advantage of incorporating the relative survival framework in mediation analysis is that it allows evaluating the impact of interventions on all-cause survival differences after targeting cancer-specific differences. This is particularly important as it isolates cancer-related differences, and it may be easier to study compared to looking at both cancer-related and other pathways. For example, through the relative survival framework is possible to estimate the potential gains in survival for those with the worst prognosis if we could eliminate stage (or treatment) differences between socioeconomic groups and improve cancer-specific survival while keeping background survival differences constant.

 

STATISTICS Seminars Series 2025-11-19: Chuang Zhang

Speaker Chuang Zhang, Beijing Wuzi University, China

Topic The Trends of AI for Meteorology

Abstract The application of AI in meteorology has become a crucial development trend. AI technologies, such as deep learning and large-scale foundation models, are rebuilding the value chain of meteorological data, covering aspects from data acquisition, cleaning, and fusion to modeling and result expression. For example, Google's AI-powered weather forecasting model simplifies meteorological forecasts, improving its efficiency and accuracy. In China, the meteorological department has developed a series of AI-based meteorological application models, significantly enhancing the level of accurate forecasting and business efficiency. However, current AI models in meteorology also have limitations. They rely heavily on historical data, and there are concerns about their effectiveness in predicting future weather patterns that may differ significantly from past patterns. Additionally, the interpretability of AI models is limited, and they struggle to identify causal relationships. The report will focus on introducing the trends in AI-empowered meteorological forecasting and the challenges associated with them. It will integrate the trend of fusion between traditional numerical models and AI, with an emphasis on describing the key technological innovations that need to be breakthroughs in the future.

 

STATISTICS Seminars Series 2025-11-26: Chamika Porage

Speaker Chamika Porage, Department of Statistics Uppsala University

Opponent Mattias Nordin, Department of Statistics, Uppsala University

Topic Evaluating model misspecification in prognostic score based average treatment effect estimation

Abstract Accurate estimation of causal effects in observational studies requires methods to account for confounding, particularly when models are subject to misspecification. Prognostic scores, which are related to outcome regression, have been introduced as an alternative to propensity scores for the estimation of the treatment effect. This study examines the performance of average treatment effect estimators based on prognostic scores and full prognostic scores (FPGS). We use various modeling approaches, including regression imputation and matching using both parametric and non-parametric techniques. Through simulation studies, we assess how model misspecification, sample size, and choice of estimator impact bias, standard error, and mean squared error of the estimator. Our findings indicate that, under correct model specification, parametric regression imputation estimators based on ordinary least squares outcome regression produce lower bias and mean squared error than regression imputation estimators based on random forest regression. In contrast, parametric regression imputation estimators exhibit a substantial increase in bias and mean squared error when the outcome regression is misspecified. For matching estimators, performance depends critically on the choice of adjustment score, particularly in the presence of heterogeneous treatment effects, where prognostic score-based matching may remain biased even when the outcome regression is correctly specified and FPGS-based matching achieves lower bias and improved mean squared error. While parametric methods perform well under correct specification, non-parametric approaches provide some flexibility against misspecification but require larger samples to achieve stability. When comparing prognostic score and FPGS-based methods, the results suggest that FPGS-based estimators may offer advantages in most cases, for example, when using regression imputation and matching estimators in the presence of heterogeneous treatment effects. To demonstrate the investigated estimators, we perform analysis in an empirical study, using data from the 2017-2018 National Health and Nutrition Examination Survey (NHANES) to investigate the effect of smoking on blood lead levels by comparing results from a correctly specified model and a misspecified model.

 

 

FOLLOW UPPSALA UNIVERSITY ON

Uppsala University on Facebook
Uppsala University on Instagram
Uppsala University on Youtube
Uppsala University on Linkedin