Previous seminars 2026

Spring

 

Review Seminar 2026-01-14: Jakob Torgander

Speaker Jakob Torgander, Department of Statistics, Uppsala University

Opponent Mattias Villani, Department of Statistics, Stockholm University

Topic SimplexMCMC: An Auxiliary-Variable Framework for Mixed Discrete–Continuous Posterior Distributions

Abstract This paper introduces SimplexMCMC, a Markov chain Monte Carlo (MCMC) framework for sampling from mixed continuous-discrete posterior distributions. Discrete sampling is reformulated as sampling from the interior of the probability simplex, leading to an auxiliary-variable MCMC construction that admits the target discrete distribution as its stationary distribution. The framework is extended to Bayesian models with mixed discrete and continuous parameters and admits an HMC-based instantiation that makes use of the Gumbel–Softmax distribution. This instantiation, referred to as Relaxed Discrete Hamiltonian Monte Carlo (RD-HMC), provides a practical algorithm for approximate inference in mixed continuous-discrete models. Correctness of the SimplexMCMC framework is established through theoretical convergence results, and the practical behaviour of RD-HMC is assessed empirically on simulated and real data.

 

STATISTICS Review Seminar 2026-02-11: Väinö Yrjänäinen

Speaker Väinö Yrjänäinen, Department of Statistics Uppsala University

Opponent Pierre Nyquist, Department of Mathematical Sciences, Chalmers University of Technology and Gothenburg University

Abstract Data accuracy is crucial for reliable research, accurate decision-making, and high-performance machine learning. However, maintaining data accuracy, especially at scale, is complicated and difficult to verify. By integrating concepts from software engineering, statistical quality control, and branching process theory, we formalize an iterative data curation framework that scales well to large data sets. We go on to prove that the proposed approach asymptotically eliminates all errors in the data with probability one. Additionally, we provide theoretical guarantees that data accuracy tests speed up error reduction. We corroborate these results through simulations on text and tabular data, and a real-world application to the Swedish Parliamentary Corpus, demonstrating the framework’s effectiveness in preserving high-accuracy historical records at scale.

 

STATISTICS Seminars Series 2026-03-11: Tilman Bretschneider

Speaker Tilman Bretschneider, Department of Economics, Lund University

Topic Factor-based imputation of missing values using cross-section averages

Abstract There is a fast-growing literature developing factor-based imputation methods of missing values in panel datasets. These methods typically estimate factors by applying principal component analysis to an observed subset of the data. In this paper, we propose a new method that uses cross section averages as factor estimates, obtains loadings by linear projections, and imputes missing values with their inner product. The main motivation behind our approach is its straightforward implementation and flexibility to accommodate general missing patterns. We derive the asymptotic properties of estimated factors, loadings, and imputed values while allowing the true number of factors to be overspecified. We showcase the good small sample properties and performance in a Monte Carlo study and in an imputation exercise with macroeconomic data. Finally, we illustrate how our method can also be used to construct monthly state-level GDP indexes from monthly macroeconomic data for a panel of U.S. states.

 

STATISTICS Industry Seminar 202603-31: Andreas Dahlström

Speaker Andreas Dahlström, Thalius AI

Topic Editable & Explainable Embeddings: Sliding Search, Taste-Vectors, and Early Semantic Arithmetic

Abstract Thalius is a deep-tech AI startup building methods for editable and explainable embeddings. In this session, we will demo a new interaction paradigm for semantic retrieval: sliding search, where users move gradually through an embedding space between arbitrary concepts instead of issuing discrete queries. We’ll show how this enables controllable product discovery and matching via learned “tastes” (preference vectors) that can be combined and tuned. Finally, I’ll share early results on semantic arithmetic: constructing, applying, and validating transformation directions to add, remove, or replace attributes in embeddings in a more predictable way. The goal is practical: make embedding systems that are not just accurate, but steerable, inspectable, and easier to debug.

 

STATISTICS Seminars Series 2026-04-15: Yingfu Xie

Speaker Yingfu Xie, Swedish Financial Supervisory Authority (Finansinspektionen)

Topic Time series forecast with statistical and Machine Learning models: a practitioner’s experience

Abstract In this talk, we go through the basic RegARIMA models for seasonal adjustment and forecasting, as well as a couple of ML models (LSTM and VAE) for time series. The talk will not be theoretical, but practical, and we will share some experiences participating in the European Statistics Awards Nowcast competition and possible issues for future research.

 

STATISTICS Seminars Series 2026-05-13: Kristofer Månsson

Speaker Kristofer Månsson, Department of Economics and Statistics, Linnaeus University

Topic Nonlinear forecasting with many predictors using mixed data sampling kernel ridge regression models

Abstract Policy institutes such as central banks need accurate forecasts of key measures of economic activity to design stabilization policies that reduce the severity of economic fluctuations. Therefore, this paper develops a kernel ridge regression estimator in a mixed data sampling framework. Kernel ridge regression can handle many predictors with a nonlinear relationship to the target variable. Consequently, it has potential to improve the currently used principal component-based methods when the economic data follow a nonlinear factor structure. In a Monte Carlo study, we show that the kernel ridge regression approach is superior in terms of mean square error and is more robust than principal component-based methods to different nonlinear data generating processes. By using a dataset consisting of 24 economic indicators, we forecast Swedish gross domestic production. The results confirm the superiority of the kernel ridge regression approach. Therefore, we suggest that policy institutes consider the use of kernel-based approaches when forecasting key measures of economic activity.

FOLLOW UPPSALA UNIVERSITY ON

Uppsala University on Facebook
Uppsala University on Instagram
Uppsala University on Youtube
Uppsala University on Linkedin