<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<atom:link href="http://jmlr.org/jmlr.xml" rel="self" type="application/rss+xml" />
<link>http://www.jmlr.org</link>
<title>JMLR</title>
<description>Journal of Machine Learning Research</description>





<item>
<title>
Towards Understanding Gradient Flow Dynamics of Homogeneous Neural Networks Beyond the Origin
</title>
<link>
http://jmlr.org/papers/v26/25-1089.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-1089/25-1089.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Akshay Kumar, Jarvis Haupt</author>
<description>
Recent works exploring the training dynamics of homogeneous neural network weights under gradient flow with small initialization have established that in the early stages of training, the weights remain small and near the origin, but converge in direction. Building on this, the current paper studies the gradient flow dynamics of homogeneous neural networks with locally Lipschitz gradients, after they escape the origin. Insights gained from this analysis are used to characterize the first saddle point encountered by gradient flow after escaping the origin. Also, it is shown that for homogeneous feed-forward neural networks, under certain conditions, the sparsity structure emerging among the weights before the escape is preserved after escaping the origin and until reaching the next saddle point.
</description>
</item>

<item>
<title>
Optimal Complexity in Byzantine-Robust Distributed Stochastic Optimization with Data Heterogeneity
</title>
<link>
http://jmlr.org/papers/v26/25-0613.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-0613/25-0613.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Qiankun Shi, Jie Peng, Kun Yuan, Xiao Wang, Qing Ling</author>
<description>
In this paper, we establish tight lower bounds for Byzantine-robust distributed first-order stochastic methods in both strongly convex and non-convex stochastic optimization. We reveal that when the distributed nodes have heterogeneous data, the convergence error comprises two components: a non-vanishing Byzantine error and a vanishing optimization error. We establish the lower bounds on the Byzantine error and on the minimum number of queries to a stochastic gradient oracle for achieving an arbitrarily small optimization error. Nevertheless, we also identify significant discrepancies between our established lower bounds and the existing upper bounds. To fill this gap, we leverage the techniques of Nesterov&#39;s acceleration and variance reduction to develop novel Byzantine-robust distributed stochastic optimization methods that provably match these lower bounds, up to at most logarithmic factors, implying that our established lower bounds are tight.
</description>
</item>

<item>
<title>
Towards Unified Native Spaces in Kernel Methods
</title>
<link>
http://jmlr.org/papers/v26/25-0022.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-0022/25-0022.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Xavier Emery, Emilio Porcu, Moreno Bevilacqua</author>
<description>
There exists a plethora of parametric models for positive definite kernels in Euclidean spaces, and their use is ubiquitous in statistics, machine learning, numerical analysis, and approximation theory. Usually, the kernel parameters index certain features of an associated process. Amongst those features, smoothness (in the sense of Sobolev spaces, mean square differentiability, and fractal dimensions), compact or global supports, and negative dependencies (hole effects) are of interest to several theoretical and applied disciplines. This paper unifies a wealth of well-known kernels into a single parametric class that encompasses them as special cases, attained either by exact parameterization or through parametric asymptotics. We furthermore find parametric restrictions under which we can characterize the Sobolev space that is norm equivalent to the RKHS associated with the new kernel. As a by-product, we infer the Sobolev spaces that are associated with existing classes of kernels. We illustrate the main properties of the new class, show how this class can switch from compact to global supports, and provide special cases for which the kernel attains negative values over nontrivial intervals. Hence, the proposed class of kernel is the reproducing kernel of a Hilbert space that contains many special cases, including the celebrated Matérn and Wendland kernels, as well as their aliases with hole effects.
</description>
</item>

<item>
<title>
TorchCP: A Python Library for Conformal Prediction
</title>
<link>
http://jmlr.org/papers/v26/24-2141.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-2141/24-2141.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jianguo Huang, Jianqing Song, Xuanning Zhou, Bingyi Jing, Hongxin Wei</author>
<description>
Conformal prediction (CP) is a powerful statistical framework that generates prediction intervals or sets with guaranteed coverage probability. While CP algorithms have evolved beyond traditional classifiers and regressors to sophisticated deep learning models like deep neural networks (DNNs), graph neural networks (GNNs), and large language models (LLMs), existing CP libraries often lack the model support and scalability for large-scale deep learning (DL) scenarios. This paper introduces TorchCP, a PyTorch-native library designed to integrate state-of-the-art CP algorithms into DL techniques, including DNN-based classifiers/regressors, GNNs, and LLMs. Released under the LGPL-3.0 license, TorchCP comprises about 16k lines of code, validated with 100% unit test coverage and detailed documentation. Notably, TorchCP enables CP-specific training algorithms, online prediction, and GPU-accelerated batch processing, achieving up to 90% reduction in inference time on large datasets. With its low-coupling design, comprehensive suite of advanced methods, and full GPU scalability, TorchCP empowers researchers and practitioners to enhance uncertainty quantification across cutting-edge applications.
</description>
</item>

<item>
<title>
Hopfield-Fenchel-Young Networks: A Unified Framework for Associative Memory Retrieval
</title>
<link>
http://jmlr.org/papers/v26/24-1961.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1961/24-1961.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Saul Santos, Vlad Niculae, Daniel McNamee, Andre F.T. Martins</author>
<description>
Associative memory models, such as Hopfield networks and their modern variants, have garnered renewed interest due to advancements in memory capacity and connections with self-attention in transformers. 
In this work, we introduce a unified framework-Hopfield-Fenchel-Young networks-which generalizes these models to a broader family of energy functions.  Our energies are formulated as the difference between two Fenchel-Young losses: one, parameterized by a generalized entropy, defines the Hopfield scoring mechanism, while the other applies a post-transformation to the Hopfield output. 
By utilizing Tsallis and norm entropies, we derive end-to-end differentiable update rules that enable sparse transformations, uncovering new connections between loss margins, sparsity, and exact retrieval of single memory patterns. 
We further extend this framework to structured Hopfield networks using the SparseMAP transformation, allowing the retrieval of  pattern associations rather than a single pattern. 
Our framework unifies and extends traditional and modern Hopfield networks and provides an energy minimization perspective for widely used post-transformations like $\ell_2$-normalization and layer normalization-all through suitable choices of Fenchel-Young losses and by using convex analysis as a building block. 
Finally, we validate our Hopfield-Fenchel-Young networks on diverse memory recall tasks, including free and sequential recall. Experiments on simulated data, image retrieval, multiple instance learning, and text rationalization demonstrate the effectiveness of our approach.
</description>
</item>

<item>
<title>
Identifiability of Causal Graphs under Non-Additive Conditionally Parametric Causal Models
</title>
<link>
http://jmlr.org/papers/v26/24-1662.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1662/24-1662.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Juraj Bodik, Valérie Chavez-Demoulin</author>
<description>
Existing approaches to causal discovery often rely on restrictive modeling assumptions that limit their applicability in real-world settings, particularly when data are heavy-tailed or contain a mixture of discrete and continuous variables. Identifiability of causal graphs has been established under several structural models, including linear non-Gaussian models, post-nonlinear models, and location-scale models. However, these frameworks may not capture the diversity of distributions observed in practice. To address this, we introduce Conditionally Parametric Causal Models (CPCM), a flexible class of models where the conditional distribution of the effect, given its cause, belongs to a known parametric family such as Gaussian, Poisson, Gamma, or Pareto. These models are adaptable to a wide range of practical situations, where the cause influences not only the mean but also the variance or tail behavior of the effect. We demonstrate the identifiability of CPCM by leveraging the concept of sufficient statistics. Furthermore, we propose an algorithm for estimating the causal structure from random samples drawn from CPCM. We evaluate the empirical properties of our methodology on various datasets, demonstrating state-of-the-art performance across multiple benchmarks.
</description>
</item>

<item>
<title>
Fundamental Limits of Membership Inference Attacks on Machine Learning Models
</title>
<link>
http://jmlr.org/papers/v26/24-1515.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1515/24-1515.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Eric Aubinais, Elisabeth Gassiat, Pablo Piantanida</author>
<description>
Membership inference attacks (MIA) can reveal whether a particular data point was part of the training dataset, potentially exposing sensitive information about individuals. This article provides theoretical guarantees by exploring the fundamental statistical limitations associated with MIAs on machine learning models at large. More precisely, we first derive the  statistical quantity that governs the effectiveness and success of such attacks. We then theoretically prove that in a  non-linear regression setting with overfitting learning procedures, attacks may have a high probability of success. Finally, we investigate several situations for which we provide bounds on this quantity of interest. Interestingly, our findings indicate that discretizing the data might enhance the learning procedure&#39;s security. Specifically, it is demonstrated to be limited by a constant, which quantifies the diversity of the underlying data distribution. We illustrate those results through simple simulations.
</description>
</item>

<item>
<title>
On the Robustness of Kernel Goodness-of-Fit Tests
</title>
<link>
http://jmlr.org/papers/v26/24-1365.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1365/24-1365.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Xing Liu, François-Xavier Briol</author>
<description>
Goodness-of-fit testing is often criticized for its lack of practical relevance: since &#34;all models are wrong&#34;, the null hypothesis that the data conform to our model is ultimately always rejected as the sample size grows. Despite this, probabilistic models are still used extensively, raising the more pertinent question of whether the model is good enough for the task at hand. This question can be formalized as a robust goodness-of-fit testing problem by asking whether the data were generated from a distribution that is a mild perturbation of the model. In this paper, we show that existing kernel goodness-of-fit tests are not robust under common notions of robustness including both qualitative and quantitative robustness. We further show that robustification techniques using tilted kernels, while effective in the  parameter estimation literature, are not sufficient to ensure both types of robustness in the testing setting. To address this, we propose the first robust kernel goodness-of-fit test, which resolves this open problem by using kernel Stein discrepancy (KSD) balls. This framework encompasses many well-known perturbation models, such as Huber&#39;s contamination and density-band models.
</description>
</item>

<item>
<title>
Efficient Online Prediction for High-Dimensional Time Series via Joint Tensor Tucker Decomposition
</title>
<link>
http://jmlr.org/papers/v26/24-1229.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1229/24-1229.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Zhenting Luan, Defeng Sun, Haoning Wang, Liping Zhang</author>
<description>
Real-time prediction plays a vital role in various control systems, such as traffic congestion control and wireless channel resource allocation. In these scenarios, the predictor usually needs to track the evolution of the latent statistical patterns in the modern high-dimensional streaming time series continuously and quickly, which presents new challenges for traditional prediction methods. This paper is the first to propose a novel online algorithm (TOPA) based on tensor factorization to predict streaming tensor time series. The proposed algorithm TOPA updates the predictor in a low-complexity online manner to adapt to the time-evolving data. Additionally, an automatically adaptive version of the algorithm (TOPA-AAW) is presented to mitigate the negative impact of stale data. Simulation results demonstrate that our proposed methods achieve prediction accuracy similar to that of conventional offline tensor prediction methods, while being much faster than them during long-term online prediction. Therefore, TOPA-AAW is an effective and efficient solution method for online prediction of streaming tensor time series.
</description>
</item>

<item>
<title>
Fast Computation of Superquantile-Constrained Optimization Through Implicit Scenario Reduction
</title>
<link>
http://jmlr.org/papers/v26/24-0752.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0752/24-0752.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jake Roth, Ying Cui</author>
<description>
Superquantiles have recently gained significant interest as a risk-aware metric for addressing fairness and distribution shifts in statistical learning and decision making problems. This paper introduces a fast, scalable and robust second-order computational framework to solve large-scale optimization problems with superquantile-based constraints. Unlike empirical risk minimization, superquantile-based optimization requires ranking random functions evaluated across all scenarios to compute the tail conditional expectation. While this tail-based feature might seem computationally unfriendly, it provides an advantageous setting for a semismooth-Newton-based augmented Lagrangian method. The superquantile operator effectively reduces the dimensions of the Newton systems since the tail expectation involves considerably fewer scenarios. Notably, the extra cost of obtaining relevant second-order information and performing matrix inversions is often comparable to, and sometimes even less than, the effort required for gradient computation. Our developed solver is particularly effective when the number of scenarios substantially exceeds the number of decision variables. In synthetic problems with linear and convex diagonal quadratic objectives, numerical experiments demonstrate that our method outperforms existing approaches by a large margin: It achieves speeds more than 750 times faster for linear and quadratic objectives than the alternating direction method of multipliers as implemented by OSQP for computing low-accuracy solutions. Additionally, it is up to 25 times faster for linear objectives and 70 times faster for quadratic objectives than the commercial solver Gurobi, and 20 times faster for linear objectives and 30 times faster for quadratic objectives than the Portfolio Safeguard optimization suite for high-accuracy solution computations. For the quantile regression problem involving over 30 million scenarios, our method computes solution paths up to 20 times faster than Gurobi. The Julia implementation of the solver is available at https://github.com/jacob-roth/superquantile-opt.
</description>
</item>

<item>
<title>
Collaborative likelihood-ratio estimation over graphs
</title>
<link>
http://jmlr.org/papers/v26/24-0565.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0565/24-0565.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Alejandro de la Concha, Nicolas Vayatis, Argyris Kalogeratos</author>
<description>
This paper introduces the Collaborative Likelihood-ratio Estimation problem, which is relevant for applications involving multiple statistical estimation tasks that can be mapped to the nodes of a fixed graph expressing pairwise task similarity. Each graph node $v$ observes i.i.d data from two unknown node-specific pdfs, $p_{v}$ and $q_{v}$, and the goal is to estimate the likelihood-ratios (or density-ratios), $r_{v}(x)=\frac{q_{v}(x)}{p_{v}(x)}$, for all $v$. Our contribution is multifold: we present a non-parametric collaborative framework that leverages the graph structure of the problem to solve the tasks more efficiently; we present a concrete method that we call Graph-based Relative Unconstrained Least-Squares Importance Fitting (GRULSIF) along with an efficient implementation; we derive convergence rates that highlight the role of the main variables of the problem. Our theoretical results explicit the conditions under which the collaborative estimation leads to performance gains compared to solving each estimation task independently. Finally, in a series of experiments, we demonstrate that the joint likelihood-ratio estimation of GRULSIF at all graph nodes is more accurate compared to state-of-the-art methods that operate independently at each node, and we verify that the behavior of GRULSIF is in agreement with our theoretical analysis.
</description>
</item>

<item>
<title>
On the Utility of Equal Batch Sizes for Inference in  Stochastic Gradient Descent
</title>
<link>
http://jmlr.org/papers/v26/24-0094.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0094/24-0094.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Rahul Singh, Abhinek Shukla, Dootika Vats</author>
<description>
Stochastic gradient descent (SGD) is an estimation tool for large data employed in machine learning and statistics. Due to the Markovian nature of the SGD process, inference is a challenging problem.  An underlying asymptotic normality of the averaged SGD (ASGD) estimator allows for the construction of a batch-means estimator of the asymptotic covariance matrix. Instead of the usual increasing batch-size strategy, we propose a memory efficient equal batch-size strategy and show that under mild conditions, the batch-means estimator is consistent. A key feature of the proposed batching technique is that it allows for bias-correction of the variance, at no additional cost to memory. Further, since joint inference for large dimensional problems may be undesirable, we present marginal-friendly simultaneous confidence intervals, and show through an example on how covariance estimators of ASGD can be employed for improved predictions.
</description>
</item>

<item>
<title>
Differentially Private Bootstrap: New Privacy Analysis and Inference Strategies
</title>
<link>
http://jmlr.org/papers/v26/23-0514.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0514/23-0514.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Zhanyu Wang, Guang Cheng, Jordan Awan</author>
<description>
Differentially private (DP) mechanisms protect individual-level information by introducing randomness into the statistical analysis procedure. Despite the availability of numerous DP tools, there remains a lack of general techniques for conducting statistical inference under DP. We examine a DP bootstrap procedure that releases multiple private bootstrap estimates to infer the sampling distribution and construct confidence intervals (CIs). Our privacy analysis presents new results on the privacy cost of a single DP bootstrap estimate, applicable to any DP mechanism, and identifies some misapplications of the bootstrap in the existing literature. For the composition of the DP bootstrap, we present a numerical method to compute the exact privacy cost of releasing multiple DP bootstrap estimates, and using the Gaussian-DP (GDP) framework (Dong et al., 2022) we show that the release of $B$ DP bootstrap estimates from mechanisms satisfying $(\mu/\sqrt{(2-2/\mathrm{e})B})$-GDP asymptotically satisfies $\mu$-GDP as $B$ goes to infinity. Then, we perform private statistical inference by post-processing the DP bootstrap estimates. We prove that our point estimates are consistent, our standard CIs are asymptotically valid, and both enjoy optimal convergence rates. To further improve the finite performance, we use deconvolution with DP bootstrap estimates to accurately infer the sampling distribution. We derive CIs for tasks such as population mean estimation, logistic regression, and quantile regression, and we compare them to existing methods using simulations and real-world experiments on 2016 Canada Census data. Our private CIs achieve the nominal coverage level and offer the first approach to private inference for quantile regression.
</description>
</item>

<item>
<title>
Convergence and Sample Complexity of Natural Policy Gradient Primal-Dual Methods for Constrained MDPs
</title>
<link>
http://jmlr.org/papers/v26/22-0622.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0622/22-0622.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Dongsheng Ding, Kaiqing Zhang, Jiali Duan, Tamer Basar, Mihailo R. Jovanovic</author>
<description>
We study the sequential decision making problem of  maximizing the expected total reward while satisfying a constraint on the expected total utility. We employ the natural policy gradient method to solve the discounted infinite-horizon optimal control problem for Constrained Markov Decision Processes (constrained MDPs). Specifically, we propose a new Natural Policy Gradient Primal-Dual (NPG-PD) method that updates the primal variable via natural policy gradient ascent and the dual variable via projected subgradient descent. Although the underlying maximization involves a nonconcave objective function and a nonconvex constraint set, under the softmax policy parametrization, we prove that our method achieves global convergence with sublinear rates regarding both the optimality gap and the constraint violation. Such convergence is independent of the size of the state-action space, i.e., it is~dimension-free. Furthermore, for log-linear and general smooth policy parametrizations, we establish sublinear convergence rates up to a function approximation error caused by restricted policy parametrization. We also provide convergence and finite-sample complexity guarantees for two sample-based NPG-PD algorithms. We use a set of computational experiments to showcase the effectiveness of our approach.
</description>
</item>

<item>
<title>
Differentially Private Multivariate Medians
</title>
<link>
http://jmlr.org/papers/v26/25-0763.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-0763/25-0763.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Kelly Ramsay, Aukosh Jagannath, Shoja&#39;eddin Chenouri</author>
<description>
Statistical tools which satisfy rigorous privacy guarantees are necessary for modern data analysis. It is well-known that robustness against contamination is linked to differential privacy. Despite this fact, using multivariate medians for differentially private and robust multivariate location estimation has not been systematically studied. We develop novel finite-sample performance guarantees for differentially private multivariate depth-based medians, which are essentially sharp. Our results cover commonly used depth functions, such as the halfspace (or Tukey) depth, spatial depth, and the integrated dual depth. We show that under Cauchy marginals, the cost of heavy-tailed location estimation outweighs the cost of privacy. We demonstrate our results numerically using a Gaussian contamination model in dimensions up to d = 100, and compare them to a state-of-the-art private mean estimation algorithm. As a by-product of our investigation, we prove concentration inequalities for the output of the exponential mechanism about the maximizer of the population objective function. This bound applies to objective functions that satisfy a mild regularity condition.
</description>
</item>

<item>
<title>
VFOSA: Variance-Reduced Fast Operator Splitting Algorithms for Generalized Equations
</title>
<link>
http://jmlr.org/papers/v26/25-0500.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-0500/25-0500.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Quoc Tran-Dinh</author>
<description>
We develop two Variance-reduced Fast Operator Splitting Algorithms (VFOSA) to approximate solutions for a class of generalized equations, covering fundamental problems such as minimization, minimax problems, and variational inequalities as special cases. Our approach integrates recent advances in accelerated operator splitting and fixed-point methods, co-hypomonotonicity structure, and variance reduction techniques. First, we introduce a class of variance-reduced estimators and establish their variance-reduction bounds. This class includes both unbiased and biased instances and comprises common estimators as special cases, including SVRG, SAGA, SARAH, and Hybrid-SGD. Second, we design a novel accelerated variance-reduced forward-backward splitting (FBS) method using these estimators to solve generalized equations in both finite-sum and expectation settings. Our algorithm achieves both O(1/k^2) and o(1/k^2) convergence rates on the expected squared norm E[ ||G_{\lambda}x^k||^2] of the FBS residual G_{\lambda}, where k is the iteration counter. Additionally, we establish almost sure convergence rates and the almost sure convergence of iterates to a
solution of the underlying generalized equation. Unlike existing stochastic operator splitting algorithms, our methods accommodate co-hypomonotone operators, which can include nonmonotone problems arising in recent applications. Third, we specify our method for each concrete estimator mentioned above and derive the corresponding oracle complexity, demonstrating that these variants achieve the best-known oracle complexity bounds without requiring additional enhancement techniques. Fourth, we develop a variance-reduced fast backward-forward splitting (BFS) method, which attains similar convergence results and oracle complexity bounds as our FBS-based algorithm. Finally, we validate our results through numerical experiments and compare their performance with existing methods.
</description>
</item>

<item>
<title>
Scaling Capability in Token Space: An Analysis of Large Vision Language Model
</title>
<link>
http://jmlr.org/papers/v26/24-2243.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-2243/24-2243.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Tenghui Li, Guoxu Zhou, Xuyang Zhao, Qibin Zhao</author>
<description>
Large language models have demonstrated predictable scaling behaviors with respect to model parameters and training data.
This study investigates whether a similar scaling relationship exist for vision-language models with respect to the number of vision tokens.
A mathematical framework is developed to characterize a relationship between vision token number and the expected divergence of distance between vision-referencing sequences.
The theoretical analysis reveals two distinct scaling regimes: sublinear scaling for less vision tokens and linear scaling for more vision tokens.
This aligns with model performance relationships of the form \(S(n) \approx c / n^{\alpha(n)}\), where the scaling exponent relates to the correlation structure between vision token representations.
Empirical validations across multiple vision-language benchmarks show that model performance matches the prediction from scaling relationship.
The findings contribute to understanding vision token scaling in transformers through a theoretical framework that complements empirical observations.
</description>
</item>

<item>
<title>
Minimax Optimal Two-Sample Testing under Local Differential Privacy
</title>
<link>
http://jmlr.org/papers/v26/24-2016.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-2016/24-2016.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jongmin Mun, Seungwoo Kwak, Ilmun Kim</author>
<description>
We explore the trade-off between privacy and statistical utility in private two-sample testing under local differential privacy (LDP) for both multinomial and continuous data. We begin with the multinomial case, where we introduce private permutation tests using practical privacy mechanisms such as Laplace, discrete Laplace, and Google&#39;s RAPPOR. We then extend this approach to continuous data via binning and study its uniform separation under LDP over H\&#34;{o}lder and Besov smoothness classes. The proposed tests for both discrete and continuous cases rigorously control type I error for any finite sample size, strictly adhere to LDP constraints, and achieve minimax optimality under LDP. The attained minimax rates reveal inherent privacy-utility trade-offs that are unavoidable in private testing. To address scenarios with unknown smoothness parameters in density testing, we propose a Bonferroni-type adaptive test that ensures robust performance without prior knowledge of the smoothness parameters. We validate our theoretical findings with extensive numerical experiments and demonstrate the practical relevance and effectiveness of our proposed methods.
</description>
</item>

<item>
<title>
Jackpot: Approximating Uncertainty Domains with Adversarial Manifolds
</title>
<link>
http://jmlr.org/papers/v26/24-1769.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1769/24-1769.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Nathanaël Munier, Emmanuel Soubies, Pierre Weiss</author>
<description>
Given a forward mapping Φ : R^N → R^M and a point x* ∈ R^N , the region {x ∈ R^N , ||Φ(x) − Φ(x*)|| ≤ ε}, where ε ≥ 0 is a perturbation amplitude, represents the set of all possible inputs x that could have produced the measurement Φ(x*) within an acceptable error margin. This set is related to uncertainty analysis, a key challenge in inverse problems. In this work, we develop a numerical algorithm called Jackpot (Jacobian Kernel Projection Optimization) which approximates this set with a low-dimensional adversarial manifold. The proposed algorithm leverages automatic differentation, allowing it to handle complex, high dimensional mappings such as those found when dealing with dynamical systems or neural networks. We demonstrate the effectiveness of our algorithm on various challenging large-scale, non-linear problems including parameter identification in dynamical systems and blind image deblurring.
</description>
</item>

<item>
<title>
An Asymptotically Optimal Coordinate Descent Algorithm for Learning Bayesian Networks from Gaussian Models
</title>
<link>
http://jmlr.org/papers/v26/24-1657.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1657/24-1657.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Tong Xu, Simge Küçükyavuz, Ali Shojaie, Armeen Taeb</author>
<description>
This paper studies the problem of learning Bayesian networks from continuous observational data, generated according to a linear Gaussian structural equation model. We consider an $\ell_0$-penalized maximum likelihood estimator for this problem, which is known to have favorable statistical properties but is computationally challenging to solve, especially for medium-sized Bayesian networks. We propose a new coordinate descent algorithm to approximate this estimator and prove several remarkable properties of our procedure: The algorithm converges to a coordinate-wise minimum, and despite the non-convexity of the loss function, as the sample size tends to infinity, the objective value of the coordinate descent solution converges to the optimal objective value of the $\ell_0$-penalized maximum likelihood estimator.  To the best of our knowledge, our proposal is the first coordinate descent procedure endowed with optimality guarantees in the context of learning Bayesian networks. Numerical experiments on synthetic and real data demonstrate that our coordinate descent method can obtain near-optimal solutions while being scalable.
</description>
</item>

<item>
<title>
Convergence Rates for Non-Log-Concave Sampling and Log-Partition Estimation
</title>
<link>
http://jmlr.org/papers/v26/24-1494.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1494/24-1494.pdf
</pdf>
<pubDate>2025</pubDate>
<author>David Holzmüller, Francis Bach</author>
<description>
Sampling from Gibbs distributions and computing their log-partition function are fundamental tasks in statistics, machine learning, and statistical physics. While efficient algorithms are known for log-concave densities, the worst-case non-log-concave setting necessarily suffers from the curse of dimensionality. For many numerical problems, the curse of dimensionality can be alleviated when the target function is smooth, allowing the exponent in the rate to improve linearly with the number of available derivatives. Recently, it has been shown that similarly fast convergence rates can be achieved by efficient optimization algorithms. Since optimization can be seen as the low-temperature limit of sampling from Gibbs distributions, we pose the question of whether similarly fast convergence rates can be achieved for non-log-concave sampling. We first study the information-based complexity of the sampling and log-partition estimation problems and show that the optimal rates for sampling and log-partition computation are sometimes equal and sometimes faster than for optimization. We then analyze various polynomial-time sampling algorithms, including an extension of a recent promising optimization approach, and find that they sometimes exhibit interesting behavior but no near-optimal rates. Our results also give further insights into the relation between sampling, log-partition, and optimization problems.
</description>
</item>

<item>
<title>
A Unified Framework to Enforce, Discover, and Promote Symmetry in Machine Learning
</title>
<link>
http://jmlr.org/papers/v26/24-1315.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1315/24-1315.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Samuel E. Otto, Nicholas Zolman, J. Nathan Kutz, Steven L. Brunton</author>
<description>
Symmetry is present throughout nature and continues to play an increasingly central role in machine learning. In this paper, we provide a unifying theoretical and methodological framework for incorporating Lie group symmetry into machine learning models in three ways: 1. enforcing known symmetry when training a model; 2. discovering unknown symmetries of a given model or data set; and 3. promoting symmetry during training by learning a model that breaks symmetries within a user-specified candidate group only when the data provide sufficient evidence. We show that these tasks can be cast within a common mathematical framework whose central object is the Lie derivative. We extend and unify several existing results by showing that enforcing and discovering symmetry are linear-algebraic tasks that are dual under the bilinear pairing induced by the Lie derivative. We also propose a novel way to promote symmetry by introducing a class of convex regularizers, built from the Lie derivative with a nuclear-norm relaxation, that penalizes symmetry breaking during training. We explain how these ideas can be applied to a wide range of machine learning models including basis-function regression, dynamical-systems discovery, neural networks, and neural operators acting on fields.
</description>
</item>

<item>
<title>
Infinite-dimensional Mahalanobis Distance with Applications to Kernelized Novelty Detection
</title>
<link>
http://jmlr.org/papers/v26/24-1126.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1126/24-1126.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Nikita Zozoulenko, Thomas Cass, Lukas Gonon</author>
<description>
The Mahalanobis distance is a classical tool used to measure the covariance-adjusted distance between points in $\mathbb{R}^d$. In this work, we extend the concept of Mahalanobis distance to separable Banach spaces by reinterpreting it as a Cameron-Martin norm associated with a probability measure. This approach leads to a basis-free, data-driven notion of anomaly distance through the so-called variance norm, which can naturally be estimated using empirical measures of a sample. Our framework generalizes the classical $\mathbb{R}^d$, functional $(L^2[0,1])^d$, and kernelized settings; importantly, it incorporates non-injective covariance operators. We prove that the variance norm is invariant under invertible bounded linear transformations of the data, extending previous results which are limited to unitary operators. In the Hilbert space setting, we connect the variance norm to the RKHS of the covariance operator, and establish consistency and convergence results for estimation using empirical measures with Tikhonov regularization. Using the variance norm, we introduce the notion of a kernelized nearest-neighbour Mahalanobis distance, and study some of its finite-sample concentration properties. In an empirical study on 12 real-world data sets, we demonstrate that the kernelized nearest-neighbour Mahalanobis distance outperforms the traditional kernelized Mahalanobis distance for multivariate time series novelty detection, using state-of-the-art time series kernels such as the signature, global alignment, and Volterra reservoir kernels.
</description>
</item>

<item>
<title>
Stable learning using spiking neural networks equipped with affine encoders and decoders
</title>
<link>
http://jmlr.org/papers/v26/24-0658.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0658/24-0658.pdf
</pdf>
<pubDate>2025</pubDate>
<author>A. Martina Neuman, Dominik Dold, Philipp Christian Petersen</author>
<description>
We study the learning problem associated with spiking neural networks. Specifically, we focus on spiking neural networks composed of simple spiking neurons having only positive synaptic weights, equipped with an affine encoder and decoder; we refer to these as affine spiking neural networks. These neural networks are shown to depend continuously on their parameters, which facilitates classical covering number-based generalization statements and supports stable gradient-based training. We demonstrate that the positivity of the weights enables a wide range of expressivity results, including rate-optimal approximation of smooth functions and dimension-independent approximation of Barron regular functions. In particular, we show in theory and simulations that affine spiking neural networks are capable of approximating shallow ReLU neural networks. Furthermore, we apply these affine spiking neural networks to standard machine learning benchmarks and reach competitive results (with respect to deep feedforward networks). Finally, we observe that from a generalization perspective, contrary to feedforward neural networks or previous results for general spiking neural networks, the depth has little to no adverse effect on theoretical guarantees for the generalization capabilities.
</description>
</item>

<item>
<title>
Efficient Knowledge Deletion from Trained Models Through Layer-wise Partial Machine Unlearning
</title>
<link>
http://jmlr.org/papers/v26/24-0469.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0469/24-0469.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Vinay Chakravarthi Gogineni, Esmaeil S. Nadimi</author>
<description>
Machine unlearning has garnered significant attention due to its ability to selectively erase knowledge obtained from specific training data samples in an already trained machine learning model. This capability enables data holders to adhere strictly to data protection regulations. However, existing unlearning techniques face practical constraints, often causing performance degradation, demanding brief fine-tuning post unlearning, and requiring significant storage. In response, this paper introduces a novel class of layer-wise partial machine unlearning algorithms that enable selective and controlled erasure of targeted knowledge. Of these, partial amnesiac unlearning integrates layer-wise selective pruning with the state-of-the-art amnesiac unlearning. This method selectively prunes and stores updates made to the model during training, enabling the targeted removal of specific data from the trained model. Other methods assimilates layer-wise partial-updates into label-flipping and optimization-based unlearning, thereby mitigating the adverse effects of specific knowledge deletion on model efficacy. Through a detailed experimental evaluation, we showcase the effectiveness of proposed unlearning methods. Experimental results highlight that the partial amnesiac unlearning not only preserves model efficacy but also eliminates the necessity for brief fine-tuning post unlearning, unlike conventional amnesiac unlearning. Further, employing layer-wise partial updates in label-flipping and optimization-based unlearning techniques demonstrates superiority in preserving model efficacy compared to their naive counterparts.
</description>
</item>

<item>
<title>
General Loss Functions Lead to (Approximate) Interpolation in High Dimensions
</title>
<link>
http://jmlr.org/papers/v26/23-1078.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1078/23-1078.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Kuo-Wei Lai, Vidya Muthukumar</author>
<description>
We provide a unified framework that applies to a general family of convex losses across binary and multiclass settings in the overparameterized regime to approximately characterize the implicit bias of gradient descent in closed form. Specifically, we show that the implicit bias is approximated (but not exactly equal to) the minimum-norm interpolation in high dimensions, which arises from training on the squared loss. In contrast to prior work, which was tailored to exponentially-tailed losses and used the intermediate support-vector-machine formulation, our framework directly builds on the primal-dual analysis of Ji and Telgarsky (2021), allowing us to provide new approximate equivalences for general convex losses through a novel sensitivity analysis. Our framework also recovers existing exact equivalence results for exponentially-tailed losses across binary and multiclass settings. Finally, we provide evidence for the tightness of our techniques and use our results to demonstrate the effect of certain loss functions designed for out-of-distribution problems on the closed-form solution.
</description>
</item>

<item>
<title>
Piecewise deterministic sampling with splitting schemes
</title>
<link>
http://jmlr.org/papers/v26/23-0036.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0036/23-0036.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Andrea Bertazzi, Paul Dobson, Pierre Monmarché</author>
<description>
We introduce Markov chain Monte Carlo (MCMC) algorithms based on numerical approximations of piecewise-deterministic Markov processes obtained with the framework of splitting schemes. We present unadjusted as well as adjusted algorithms, for which the asymptotic bias due to the discretisation error is removed applying a non-reversible Metropolis-Hastings filter. In a general framework we demonstrate that the unadjusted schemes have weak error of second order in the step size, while typically maintaining a computational cost of only one gradient evaluation of the negative log-target function per iteration. Focusing then on unadjusted schemes based on the Bouncy Particle and Zig-Zag samplers, we provide conditions ensuring geometric ergodicity and consider the expansion of the invariant measure in terms of the step size. We analyse the dependence of the leading term in this expansion on the refreshment rate and on the structure of the splitting scheme, giving a guideline on which structure is best. Finally, we illustrate promising results for our samplers with numerical experiments on a Bayesian imaging inverse problem and a system of interacting particles.
</description>
</item>

<item>
<title>
Hierarchical and Stochastic Crystallization Learning: Geometrically Leveraged Nonparametric Regression with Delaunay Triangulation
</title>
<link>
http://jmlr.org/papers/v26/21-0930.html
</link>
<pdf>
http://jmlr.org/papers/volume26/21-0930/21-0930.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jiaqi Gu, Guosheng Yin</author>
<description>
High-dimensionality is known to be the bottleneck for both nonparametric regression and the Delaunay triangulation. To efficiently exploit the advantage of the Delaunay triangulation in utilizing geometry information for nonparametric regression without conducting the Delaunay triangulation for the entire feature space, we develop the crystallization search for the neighbor Delaunay simplices of the target point similar to crystal growth and estimate the conditional expectation function by fitting a local linear model to the data points of the constructed Delaunay simplices. Because the shapes and volumes of Delaunay simplices are adaptive to the density of feature data points, our method selects neighbor data points more uniformly in all directions in comparison with Euclidean distance based methods and thus it is more robust to the local geometric structure of the data.  We further develop the stochastic approach to hyperparameter selection and the hierarchical crystallization learning under multimodal feature data densities, where an approximate global Delaunay triangulation is obtained by first triangulating the local centers and then constructing local Delaunay triangulations in parallel.  We study the asymptotic properties of our method and conduct numerical experiments on both synthetic and real data to demonstrate the advantages of our method over the existing ones.
</description>
</item>

<item>
<title>
Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2
</title>
<link>
http://jmlr.org/papers/v26/25-1654.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-1654/25-1654.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yuri Chervonyi, Trieu H. Trinh, Miroslav Olšák, Xiaomeng Yang, Hoang H. Nguyen, Marcelo Menegali, Junehyuk Jung, Junsu Kim, Vikas Verma, Quoc V. Le, Thang Luong</author>
<description>
We present AlphaGeometry2, a significantly improved version of AlphaGeometry introduced in Nature, 625 (7995):476, 2024, which has now surpassed an average gold medalist in solving Olympiad geometry problems. To achieve this, we first extend the original AlphaGeometry language to tackle problems involving movements of objects, and problems containing linear equations of angles, ratios, and distances. This, together with support for non-constructive problems, has markedly improved the coverage rate of the AlphaGeometry language on International Math Olympiads 2000-2024 geometry problems from 66% to 88%. The search process of AlphaGeometry2 has also been greatly improved through the use of Gemini architecture for better language modeling, and a novel knowledge-sharing mechanism that enables effective communication between search trees. Together with further enhancements to the symbolic engine and synthetic data generation, we have significantly boosted the overall solving rate of AlphaGeometry to 84% on all geometry problems over the last 25 years, compared to 54% previously. AlphaGeometry2 was also part of the system that achieved the silver-medal standard at IMO 2024 https://dpmd.ai/imo-silver. Finally, we report progress towards using AlphaGeometry2 as a part of a fully automated system that reliably solves geometry problems from natural language input. Code: https://github.com/google-deepmind/alphageometry2.
</description>
</item>

<item>
<title>
Decentralized Bilevel Optimization: A Perspective from Transient Iteration Complexity
</title>
<link>
http://jmlr.org/papers/v26/25-0677.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-0677/25-0677.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Boao Kong, Shuchen Zhu, Songtao Lu, Xinmeng Huang, Kun Yuan</author>
<description>
Stochastic bilevel optimization (SBO) is becoming increasingly essential in machine learning due to its versatility in handling nested structures. To address large-scale SBO, decentralized approaches have emerged as effective paradigms in which nodes communicate with immediate neighbors without a central server, thereby improving communication efficiency and enhancing algorithmic robustness. However, most decentralized SBO algorithms focus solely on asymptotic convergence rates, overlooking transient iteration complexity-the number of iterations required before asymptotic rates dominate, which results in limited understanding of the influence of network topology, data heterogeneity, and the nested bilevel algorithmic structures. To address this issue, this paper introduces D-SOBA, a Decentralized Stochastic One-loop Bilevel Algorithm framework. D-SOBA comprises two variants: D-SOBA-SO, which incorporates second-order Hessian and Jacobian matrices, and D-SOBA-FO, which relies entirely on first-order gradients. We provide a comprehensive non-asymptotic convergence analysis and establish the transient iteration complexity of D-SOBA. This provides the first theoretical understanding of how network topology, data heterogeneity, and nested bilevel structures influence decentralized SBO. Extensive experimental results demonstrate the efficiency and theoretical advantages of D-SOBA.
</description>
</item>

<item>
<title>
Fair Text Classification via Transferable Representations
</title>
<link>
http://jmlr.org/papers/v26/25-0485.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-0485/25-0485.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Thibaud Leteno, Michael Perrot, Charlotte Laclau, Antoine Gourru, Christophe Gravier</author>
<description>
Group fairness is a central research topic in text classification, where reaching fair treatment between sensitive groups (e.g., women and men) remains an open challenge. We propose an approach that extends the use of the Wasserstein Dependency Measure for learning unbiased neural text classifiers.
Given the challenge of distinguishing fair from unfair information in a text encoder, we draw inspiration from adversarial training by inducing independence between representations learned for the target label and those for a sensitive attribute. We further show that domain adaptation can be efficiently leveraged to remove the need for access to the sensitive attributes in the data set we cure. We provide both theoretical and empirical evidence that our approach is well-founded.
</description>
</item>

<item>
<title>
Stochastic Interior-Point Methods for Smooth Conic Optimization with Applications
</title>
<link>
http://jmlr.org/papers/v26/24-2158.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-2158/24-2158.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Chuan He, Zhanwang Deng</author>
<description>
Conic optimization plays a crucial role in many machine learning (ML) problems. However, practical algorithms for conic constrained ML problems with large datasets are often limited to specific use cases, as stochastic algorithms for general conic optimization remain underdeveloped. To fill this gap, we introduce a stochastic interior-point method (SIPM) framework for general conic optimization, along with four novel SIPM variants leveraging distinct stochastic gradient estimators. Under mild assumptions, we establish the iteration complexity of our proposed SIPMs, which, up to a polylogarithmic factor, matches the best-known results in stochastic unconstrained optimization. Finally, our numerical experiments on robust linear regression, multi-task relationship learning, and clustering data streams demonstrate the effectiveness and efficiency of our approach.
</description>
</item>

<item>
<title>
Revisiting Gradient Normalization and Clipping for Nonconvex SGD under Heavy-Tailed Noise: Necessity, Sufficiency, and Acceleration
</title>
<link>
http://jmlr.org/papers/v26/24-1991.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1991/24-1991.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Tao Sun, Xinwang Liu, Kun Yuan</author>
<description>
Gradient clipping has long been considered essential for ensuring the convergence of Stochastic Gradient Descent (SGD) in the presence of heavy-tailed gradient noise. In this paper, we revisit this belief and explore whether gradient normalization can serve as an effective alternative or complement. We prove that, under  individual smoothness assumptions, gradient normalization alone is sufficient to guarantee convergence of
 the nonconvex SGD. Moreover, when combined with clipping, it yields far better rates of convergence under   more challenging noise distributions. We provide a unifying theory describing normalization-only, clipping-only, and combined approaches. Moving forward, we investigate existing variance-reduced algorithms, establishing that, in such a setting, normalization alone is sufficient for convergence. Finally, we present an accelerated variant that under second-order smoothness improves convergence.
Our results provide theoretical insights and practical guidance for using normalization and clipping in nonconvex optimization with heavy-tailed noise.
</description>
</item>

<item>
<title>
Generalized multi-view model: Adaptive density estimation under low-rank constraints
</title>
<link>
http://jmlr.org/papers/v26/24-1729.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1729/24-1729.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Julien Chhor, Olga Klopp, Alexandre B. Tsybakov</author>
<description>
We study the problem of bivariate discrete or continuous probability density estimation under low-rank constraints. For discrete distributions, we assume that the two-dimensional array to estimate is a low-rank probability matrix. In the continuous case, we assume that the density with respect to the Lebesgue measure satisfies a generalized multi-view model, meaning that it is $\beta$-Hölder and can be decomposed as a sum of $K$ components, each of which is a product of one-dimensional functions. In both settings, we propose estimators that achieve, up to logarithmic factors, the minimax optimal convergence rates under such low-rank constraints. In the discrete case, the proposed estimator is adaptive to the rank $K$. In the continuous case, our estimator converges with the $L_1$ rate $\min((K/n)^{\beta/(2\beta+1)}, n^{-\beta/(2\beta+2)})$ up to logarithmic factors, and it is adaptive to the unknown support as well as to the smoothness $\beta$ and to the unknown number of separable components $K$. We present efficient algorithms to compute our estimators.
</description>
</item>

<item>
<title>
(De)-regularized Maximum Mean Discrepancy Gradient Flow
</title>
<link>
http://jmlr.org/papers/v26/24-1574.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1574/24-1574.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Zonghao Chen, Aratrika Mustafi, Pierre Glaser, Anna Korba, Arthur Gretton, Bharath K. Sriperumbudur</author>
<description>
We introduce a (de)-regularization of the Maximum Mean Discrepancy (DrMMD) and its Wasserstein gradient flow. Existing gradient flows that transport samples from source distribution to target distribution with only target samples, either lack tractable numerical implementation ($f$-divergence flows) or require strong assumptions and modifications, such as noise injection, to ensure convergence (Maximum Mean Discrepancy flows). In contrast, DrMMD flow can simultaneously (i) guarantee near-global convergence for a broad class of targets in both continuous and discrete time, and (ii) be implemented in closed form using only samples. The former is achieved by leveraging the connection between the DrMMD and the $\chi^2$-divergence, while the latter comes by treating DrMMD as MMD with a de-regularized kernel. Our numerical scheme employs an adaptive de-regularization schedule throughout the flow to optimally balance the trade-off between discretization errors and deviations from the $\chi^2$ regime. The potential application of the DrMMD flow is demonstrated across several numerical experiments, including a large-scale setting of training student/teacher networks.
</description>
</item>

<item>
<title>
On Probabilistic Embeddings in Optimal Dimension Reduction
</title>
<link>
http://jmlr.org/papers/v26/24-1461.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1461/24-1461.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Ryan Murray, Adam Pickarski</author>
<description>
Dimension reduction algorithms are essential in data science for tasks such as data exploration, feature selection, and denoising. However, many non-linear dimension reduction algorithms are poorly understood from a theoretical perspective. This work considers a generalized version of multidimensional scaling, which seeks to construct a map from high to low dimension which best preserves pairwise inner products or norms. We investigate the variational properties of this problem, leading to the following insights: 1) Particle-wise descent methods implemented in standard libraries can produce non-deterministic embeddings, 2) A probabilistic formulation leads to solutions with interpretable necessary conditions, and 3) The globally optimal solutions to the relaxed, probabilistic problem is only minimized by deterministic embeddings. This progression of results mirrors the classical development of optimal transportation, and in a case relating to the Gromov-Wasserstein distance actually gives explicit insight into the structure of the optimal embeddings, which are parametrically determined and discontinuous on smooth surfaces. Our results also imply that a standard computational implementation for this problem learns sub-optimal mappings, and we discuss how the embeddings learned in that context have highly misleading clustering structure, underscoring the delicate nature of solving this problem computationally.
</description>
</item>

<item>
<title>
Physics Informed Kolmogorov-Arnold Neural Networks for Dynamical Analysis via Efficient-KAN and WAV-KAN
</title>
<link>
http://jmlr.org/papers/v26/24-1278.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1278/24-1278.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Subhajit Patra, Sonali Panda, Bikram Keshari Parida, Mahima Arya, Kurt Jacobs, Denys I. Bondar, Abhijit Sen</author>
<description>
Physics-informed neural networks have proven to be a powerful tool for solving differential equations, leveraging the principles of physics to inform the learning process. However, traditional deep neural networks often face challenges in achieving high accuracy without incurring significant computational costs. In this work, we implement the Physics-Informed Kolmogorov-Arnold Neural Networks (PIKAN) through efficient-KAN and WAV-KAN, which utilize the Kolmogorov-Arnold representation theorem. PIKAN demonstrates superior performance compared to conventional deep neural networks, achieving the same level of accuracy with fewer layers and reduced computational overhead. We explore both B-spline and wavelet-based implementations of PIKAN and benchmark their performance across various ordinary and partial differential equations using unsupervised (data-free) and supervised (data-driven) techniques. For certain differential equations, the data-free approach suffices to find accurate solutions, while in more complex scenarios, the data-driven method enhances the PIKAN’s ability to converge to the correct solution. We validate our results against numerical solutions and achieve $99\%$ accuracy in most scenarios.
</description>
</item>

<item>
<title>
Graph-accelerated Markov Chain Monte Carlo using Approximate Samples
</title>
<link>
http://jmlr.org/papers/v26/24-1024.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1024/24-1024.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Leo L. Duan, Anirban Bhattacharya</author>
<description>
It has become increasingly easy nowadays to collect approximate posterior samples via fast algorithms such as variational Bayes, but concerns exist about the estimation accuracy. It is tempting to build solutions that exploit approximate samples in a canonical Markov chain Monte Carlo framework. As the dimension increases, a major barrier is that the approximate sample tends to have a low Metropolis--Hastings acceptance rate when used as a proposal. In this article, we propose a simple solution named graph-accelerated Markov Chain Monte Carlo. We build a graph with each node assigned to an approximate sample, then run Markov chain Monte Carlo with random walks over the graph. We optimize the graph edges to enforce small differences in posterior density/probability between nodes, while encouraging edges to have large distances in the parameter space. The graph allows us to accelerate a canonical Markov transition kernel through mixing with a large-jump Metropolis-Hastings step. The acceleration is easily applicable to existing Markov chain Monte Carlo algorithms. We theoretically quantify the rate of acceptance as dimension increases, and show the effects on improved mixing time. We demonstrate improved mixing performances for challenging problems, such as those involving multiple modes, non-convex density contour, or large-dimension latent variables.
</description>
</item>

<item>
<title>
Online Quantile Regression
</title>
<link>
http://jmlr.org/papers/v26/24-0589.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0589/24-0589.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yinan Shen, Dong Xia, Wen-Xin Zhou</author>
<description>
This paper addresses the challenge of integrating sequentially arriving data into the quantile regression framework, where the number of features may increase with the number of observations, the time horizon is unknown, and memory resources are limited. Unlike least squares and robust regression methods, quantile regression models different segments of the conditional distribution, thereby capturing heterogeneous relationships between predictors and responses and providing a more comprehensive view of the underlying stochastic structure. We employ stochastic sub-gradient descent to minimize the empirical check loss and analyze its statistical properties and regret behavior. Our analysis reveals a subtle interplay between updating iterates based on individual observations and on batches of observations, highlighting distinct regularity characteristics in each setting. The proposed method guarantees long-term optimal estimation performance regardless of the chosen update strategy. Our contributions extend existing literature by establishing exponential-type concentration inequalities and by achieving optimal regret and error rates that exhibit only short-term sensitivity to initialization. A key insight from our study lies in the refined statistical analysis showing that properly chosen stepsize schemes substantially mitigate the influence of initial errors on subsequent estimation and regret. This result underscores the robustness of stochastic sub-gradient descent in managing initial uncertainties and affirms its effectiveness in sequential learning settings with unknown horizons and data-dependent sample sizes. Furthermore, when the initial estimation error is well-controlled, our analysis reveals a trade-off between short-term error reduction and long-term optimality. For completeness, we also discuss the squared loss case and outline appropriate update schemes, whose analysis requires additional care. Extensive simulation studies corroborate our theoretical findings.
</description>
</item>

<item>
<title>
Statistical Inference of Random Graphs With a Surrogate Likelihood Function
</title>
<link>
http://jmlr.org/papers/v26/24-0153.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0153/24-0153.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Dingbo Wu, Fangzheng Xie</author>
<description>
Spectral estimators have been broadly applied to statistical network analysis, but they do not incorporate the likelihood information of the network sampling model. This paper proposes a novel surrogate likelihood function for statistical inference of a class of popular network models referred to as random dot product graphs. In contrast to the structurally complicated exact likelihood function, the surrogate likelihood function has a separable structure and is log-concave yet approximates the exact likelihood function well. From the frequentist perspective, we study the maximum surrogate likelihood estimator and establish the accompanying theory. We show its existence, uniqueness, large sample properties, and that it improves upon the baseline spectral estimator with a smaller sum of squared errors. Furthermore, we derive the second-order bias of the proposed estimator and gain insight into why it outperforms some of the existing estimators. A computationally convenient stochastic gradient descent algorithm is designed to find the maximum surrogate likelihood estimator in practice. From the Bayesian perspective, we establish the Bernstein--von Mises theorem of the posterior distribution with the surrogate likelihood function and show that the resulting credible sets have the correct frequentist coverage. The empirical performance of the proposed surrogate-likelihood-based methods is validated through the analyses of simulation examples and two real-world data sets.
</description>
</item>

<item>
<title>
On the Representation of Pairwise Causal Background Knowledge and Its Applications in Causal Inference
</title>
<link>
http://jmlr.org/papers/v26/23-0624.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0624/23-0624.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Zhuangyan Fang, Ruiqi Zhao, Yue Liu, Yangbo He</author>
<description>
Pairwise causal background knowledge about the existence or absence of causal edges and paths is frequently encountered in observational studies. Such constraints allow the shared directed and undirected edges in the constrained subclass of Markov equivalent DAGs to be represented as a causal maximally partially directed acyclic graph (MPDAG). In this paper, we first provide a sound and complete graphical characterization of causal MPDAGs and introduce a minimal representation of a causal MPDAG. Then, we give a unified representation for three types of pairwise causal background knowledge, including direct, ancestral and non-ancestral causal knowledge, by introducing a novel concept called direct causal clause (DCC). Using DCCs, we study the consistency and equivalence of pairwise causal background knowledge and show that any pairwise causal background knowledge set can be uniquely and equivalently decomposed into the causal MPDAG representing the refined Markov equivalence class and a minimal residual set of DCCs. Polynomial-time algorithms are also provided for checking consistency and equivalence, as well as for finding the decomposed MPDAG and the residual DCCs. Finally, with pairwise causal background knowledge, we prove a sufficient and necessary condition to identify causal effects and surprisingly find that the identifiability of causal effects only depends on the decomposed MPDAG. We also develop a local IDA-type algorithm to estimate the possible values of an unidentifiable effect. Simulations suggest that pairwise causal background knowledge can significantly improve the identifiability of causal effects.
</description>
</item>

<item>
<title>
An Augmentation Overlap Theory of Contrastive Learning
</title>
<link>
http://jmlr.org/papers/v26/22-1009.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-1009/22-1009.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Qi Zhang, Yifei Wang, Yisen Wang</author>
<description>
Recently, self-supervised contrastive learning has achieved great success on various tasks. However, its underlying working mechanism is yet unclear. In this paper, we first provide the tightest bounds based on the widely adopted assumption of conditional independence. Further, we relax the conditional independence assumption to a more practical assumption of augmentation overlap and derive the asymptotically closed bounds for the downstream performance. Our proposed augmentation overlap theory hinges on the insight that the support of different intra-class samples will become more overlapped under aggressive data augmentations, thus simply aligning the positive samples (augmented views of the same sample) could make contrastive learning cluster intra-class samples together. Moreover, from the newly derived augmentation overlap perspective, we develop an unsupervised metric for the representation evaluation of contrastive learning, which aligns well with the downstream performance almost without relying on additional modules. Code is available at https://github.com/PKU-ML/GARC.
</description>
</item>

<item>
<title>
Algorithms for ridge estimation with convergence guarantees
</title>
<link>
http://jmlr.org/papers/v26/21-0095.html
</link>
<pdf>
http://jmlr.org/papers/volume26/21-0095/21-0095.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Wanli Qiao, Wolfgang Polonik</author>
<description>
The extraction of filamentary structure from a point cloud is discussed. The filaments are modeled as ridge lines or higher dimensional ridges of an underlying density. We propose two novel algorithms, and provide theoretical guarantees for their convergences, by which we mean that the algorithms can asymptotically recover the full ridge set. We consider the new algorithms as alternatives to the Subspace Constrained Mean Shift (SCMS) algorithm for which no such theoretical guarantees are known.
</description>
</item>

<item>
<title>
Talent: A Tabular Analytics and Learning Toolbox
</title>
<link>
http://jmlr.org/papers/v26/25-0512.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-0512/25-0512.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Si-Yang Liu, Hao-Run Cai, Qi-Le Zhou, Huai-Hong Yin, Tao Zhou, Jun-Peng Jiang, Han-Jia Ye</author>
<description>
Tabular data is a prevalent source in machine learning. While classical methods have proven effective, deep learning methods for tabular data are emerging as flexible alternatives due to their capacity to uncover hidden patterns and capture complex interactions. Considering that deep tabular methods exhibit diverse design philosophies, including the ways they handle features, design learning objectives, and construct model architectures, we introduce Talent (Tabular Analytics and Learning Toolbox), a versatile toolbox for utilizing, analyzing, and comparing these methods.  Talent includes over 35 deep tabular prediction methods, offering various encoding and normalization modules, all within a unified, easily extensible interface. We demonstrate its design, application, and performance evaluation in case studies. The code is available at https://github.com/LAMDA-Tabular/TALENT.
</description>
</item>

<item>
<title>
Inferring Change Points in High-Dimensional Regression via Approximate Message Passing
</title>
<link>
http://jmlr.org/papers/v26/24-1789.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1789/24-1789.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Gabriel Arpino, Xiaoqi Liu, Julia Gontarek, Ramji Venkataramanan</author>
<description>
We consider the problem of localizing change points in a generalized linear model (GLM), a model that covers many widely studied problems in statistical learning including linear, logistic, and rectified linear regression. We propose a novel and computationally efficient approximate message passing (AMP) algorithm for estimating both the signals and the change point locations, and rigorously characterize its performance in the high-dimensional limit where the number of parameters $p$ is proportional to the number of samples $n$. This characterization is in terms of a state evolution recursion, which allows us to precisely compute performance measures such as the asymptotic Hausdorff error of our change point estimates, and allows us to tailor the algorithm to take advantage of any prior structural information of the signals and change points. Moreover, we show how our AMP iterates can be used to efficiently compute a Bayesian posterior distribution over the change point locations in the high-dimensional limit. We validate our theory via numerical experiments, and demonstrate the favorable performance of our estimators on both synthetic and real data in the settings of linear, logistic, and rectified linear regression.
</description>
</item>

<item>
<title>
Universality of Kernel Random Matrices and Kernel Regression in the Quadratic Regime
</title>
<link>
http://jmlr.org/papers/v26/24-1533.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1533/24-1533.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Parthe Pandit, Zhichao Wang, Yizhe Zhu</author>
<description>
Kernel ridge regression (KRR) is a popular class of machine learning models that has become an important tool for understanding deep learning.  Much of the focus thus far has been on studying the proportional asymptotic regime, $n \asymp d$, where $n$ is the number of training samples and $d$ is the dimension of the dataset. In the proportional regime, under certain conditions on the data distribution, the kernel random matrix involved in KRR exhibits behavior akin to that of a linear kernel. In this work, we extend the study of kernel regression to the quadratic asymptotic regime, where $n \asymp d^2$. In this regime, we demonstrate that a broad class of inner-product kernels exhibits behavior similar to a quadratic kernel. Specifically, we establish an operator norm approximation bound for the difference between the original kernel random matrix and a quadratic kernel random matrix with additional correction terms compared to the Taylor expansion of the kernel functions. The approximation works for general data distributions under a Gaussian-moment-matching assumption with a covariance structure. This new approximation is utilized to obtain a limiting spectral distribution of the original kernel matrix and characterize the precise asymptotic training and test errors for KRR in the quadratic regime when $n/d^2$ converges to a non-zero constant. The generalization errors are obtained for (i) a random teacher model, (ii) a deterministic teacher model where the weights are perfectly aligned with the covariance of the data. Under the random teacher model setting, we also verify that the generalized cross-validation (GCV) estimator can consistently estimate the generalization error in the quadratic regime for anisotropic data. Our proof techniques combine moment methods, Wick&#39;s formula, orthogonal polynomials, and resolvent analysis of random matrices with correlated entries.
</description>
</item>

<item>
<title>
Lexicographic Lipschitz Bandits: New Algorithms and a Lower Bound
</title>
<link>
http://jmlr.org/papers/v26/24-1047.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1047/24-1047.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Bo Xue, Ji Cheng, Fei Liu, Yimu Wang, Lijun Zhang, Qingfu Zhang</author>
<description>
This paper studies a multiobjective bandit problem under lexicographic ordering, wherein the learner aims to maximize $m$ objectives, each with different levels of importance. First, we introduce the local trade-off, $\lambda_*$, which depicts the trade-off between different objectives. For the case when an upper bound of $\lambda_*$ is known, i.e., $\lambda\geq\lambda_*$, we develop an algorithm that achieves a general regret bound of $\widetilde{O}(\Lambda^i(\lambda)T^{(d_z^i+1)/(d_z^i+2)})$ for the $i$-th objective, where $i\in\{1,2,\ldots,m\}$, $\Lambda^i(\lambda)=1+\lambda+\cdots+\lambda^{i-1}$, $d_z^i$ is the zooming dimension for the $i$-th objective, and $T$ is the time horizon. Next, we provide a matching lower bound for the lexicographic Lipschitz bandit problem, proving that our algorithm is optimal in terms of $\lambda_*$ and $T$. Finally, for the case where $m=2$, we remove the dependence on the knowledge about $\lambda_*$, albeit at the cost of increasing the regret bound to $\widetilde{O}(\Lambda^i(\lambda_*)T^{(3d_z^i+4)/(3d_z^i+6)})$, which remains optimal in terms of $\lambda_*$. Compared to existing work on lexicographic multi-armed bandits, our approach improves the current regret bound of $\widetilde{O}(T^{2/3})$ and extends the number of arms to infinity. Numerical experiments confirm the effectiveness of our algorithms.
</description>
</item>

<item>
<title>
On the Natural Gradient of the Evidence Lower Bound
</title>
<link>
http://jmlr.org/papers/v26/24-0606.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0606/24-0606.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Nihat Ay, Jesse van Oostrum, Adwait Datar</author>
<description>
This article studies the Fisher-Rao gradient, also referred to as the natural gradient, of the evidence lower bound (ELBO) which plays a central role in generative machine learning. It reveals that the gap between the evidence and its lower bound, the ELBO, has essentially a vanishing natural gradient within unconstrained optimization. As a result, maximization of the ELBO is equivalent to minimization of the Kullback-Leibler divergence from a target distribution, the primary objective function of learning. Building on this insight, we derive a condition under which this equivalence persists even when optimization is constrained to a model. This condition yields a geometric characterization, which we formalize through the notion of a cylindrical model.
</description>
</item>

<item>
<title>
Geometry and Stability of Supervised Learning Problems
</title>
<link>
http://jmlr.org/papers/v26/24-0322.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0322/24-0322.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Facundo Mémoli, Brantley Vose, Robert C. Williamson</author>
<description>
We introduce a notion of distance between supervised learning problems, which we call the Risk distance. This distance, inspired by optimal transport, facilitates stability results; one can quantify how seriously issues like sampling bias, noise, limited data, and approximations might change a given problem by bounding how much these modifications can move the problem under the Risk distance. With the distance established, we explore the geometry of the resulting space of supervised learning problems, providing explicit geodesics and proving that the set of classification problems is dense in a larger class of problems. We also provide two variants of the Risk distance: one that incorporates specified weights on a problem&#39;s predictors, and one that is more sensitive to the contours of a problem&#39;s risk landscape.
</description>
</item>

<item>
<title>
Understanding Deep Representation Learning via Layerwise Feature Compression and Discrimination
</title>
<link>
http://jmlr.org/papers/v26/24-0047.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0047/24-0047.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Peng Wang, Xiao Li, Can Yaras, Zhihui Zhu, Laura Balzano, Wei Hu, Qing Qu</author>
<description>
Over the past decade, deep learning has proven to be a highly effective tool for learning meaningful features from raw data. However, it remains an open question how deep networks perform hierarchical feature learning across layers. In this work, we attempt to unveil this mystery by investigating the structures of intermediate features. Motivated by our empirical findings that linear layers mimic the roles of deep layers in nonlinear networks for feature learning, we explore how deep linear networks transform input data into output by investigating the output (i.e., features) of each layer after training in the context of multi-class classification problems. Toward this goal, we first define metrics to measure within-class compression and between-class discrimination of intermediate features, respectively. Through theoretical analysis of these two metrics, we show that the evolution of features follows a simple and quantitative pattern from shallow to deep layers when the input data is nearly orthogonal and the network weights are minimum-norm, balanced, and approximately low-rank: each layer of the linear network progressively compresses within-class features at a geometric rate and discriminates between-class features at a linear rate with respect to the number of layers that data have passed through. To the best of our knowledge, this is the first quantitative characterization of feature evolution in hierarchical representations of deep linear networks. Moreover, our extensive experiments not only validate our theoretical results but also reveal a similar pattern in deep nonlinear networks, which aligns well with recent empirical studies. Finally, we demonstrate the practical value of our results in transfer learning.
</description>
</item>

<item>
<title>
Optimal Rates of Kernel Ridge Regression under Source Condition in Large Dimensions
</title>
<link>
http://jmlr.org/papers/v26/23-1679.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1679/23-1679.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Haobo Zhang, Yicheng Li, Weihao Lu, Qian Lin</author>
<description>
Motivated by studies of neural networks, particularly the neural tangent kernel theory, we investigate the large-dimensional behavior of kernel ridge regression, where the sample size satisfies $n $ is proportion to $ d^{\gamma}$ for some $\gamma &gt; 0$. Given a reproducing kernel Hilbert space $H$ associated with an inner product kernel defined on the unit sphere $S^{d}$, we assume that the true function $f_{\rho}^{*}$ belongs to the interpolation space $[H]^{s}$ for some $s&gt;0$ (source condition). We first establish the exact order (both upper and lower bounds) of the generalization error of KRR for the optimally chosen regularization parameter $\lambda$. Furthermore, we show that KRR is minimax optimal when $0&lt;s\le1$, whereas for $s&gt;1$, KRR fails to achieve minimax optimality, exhibiting the saturation effect. Our results illustrate that the convergence rate with respect to dimension $d$ varying along $\gamma$ exhibits a periodic plateau behavior, and the convergence rate with respect to sample size $n$ exhibits a multiple descent behavior. Interestingly, our work unifies several recent studies on kernel regression in the large-dimensional setting, which correspond to $s=0$ and $s=1$, respectively.
</description>
</item>

<item>
<title>
A Hybrid Weighted Nearest Neighbour Classifier for Semi-Supervised Learning
</title>
<link>
http://jmlr.org/papers/v26/23-1258.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1258/23-1258.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Stephen M. S. Lee, Mehdi Soleymani</author>
<description>
We propose a novel hybrid procedure for constructing a randomly weighted nearest neighbour classifier for semi-supervised learning. The procedure first uses the labelled learning set to predict a probability distribution of class labels for the unlabelled learning set. This turns the unlabelled set into a pseudo-labelled set, on which a sequentially weighted nearest neighbour classifier can be trained. The vote proportions calculated by this sequentially weighted nearest neighbour classifier and the standard weighted nearest neighbour classifier trained on the labelled set alone are then linearly combined to build a hybrid classifier. Our theory shows that, given a sufficiently large set of unlabelled data, the hybrid classifier has an optimal regret converging at a faster rate than that of the optimally weighted nearest neighbour classifier and hence of the optimal bagged or k-nearest neighbour classifier. We also show that the hybrid classifier can be revised by a dislabelling strategy to achieve the fastest possible rate of regret irrespective of the size of the unlabelled set, which may even be empty.  Simulation studies and real data examples are presented to support our theoretical findings and illustrate the empirical performance of the hybrid classifiers constructed using uniform weights. We also explore the effects of pseudo-labelling by hypothesized class probabilities as a supplement to our main findings.
</description>
</item>

<item>
<title>
Scalable and Adaptive Variational Bayes Methods for Hawkes Processes
</title>
<link>
http://jmlr.org/papers/v26/23-1053.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1053/23-1053.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Deborah Sulem, Vincent Rivoirard, Judith Rousseau</author>
<description>
Hawkes processes are often applied to model dependence and interaction phenomena in multivariate event data sets, such as neuronal spike trains, social interactions, and financial transactions. In the nonparametric setting, learning the temporal dependence structure of Hawkes processes is generally a computationally expensive task, all the more with Bayesian estimation methods. In particular, for multivariate nonlinear Hawkes processes, Monte-Carlo Markov Chain (MCMC) methods used to sample from the posterior distribution do not scale well to the dimension of the process. Recently, efficient algorithms targeting a mean-field variational approximation of the posterior distribution have been proposed, however, these methods do not allow to perform model selection on the graph of interactions of the Hawkes model. In this work, we propose a novel adaptive Bayesian variational method that performs model selection and can estimate a sparse graphical parameter. For the popular sigmoid Hawkes processes, we design a parallel algorithm which is scalable to high-dimensional point processes and large sequences of events. Furthermore, we unify existing variational Bayes approaches under a general nonparametric inference framework, and analyse the asymptotic properties of these methods under easily verifiable conditions on the prior, the variational class, and the nonlinear model. Finally, through an extensive set of numerical simulations, we demonstrate that our method is able to adapt to the dimensionality of the parameter of the Hawkes process, and is partially robust to certain types of model misspecification.
</description>
</item>

<item>
<title>
Biological Sequence Kernels with Guaranteed Flexibility
</title>
<link>
http://jmlr.org/papers/v26/23-0455.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0455/23-0455.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Alan N. Amin, Debora S. Marks, Eli N. Weinstein</author>
<description>
Applying machine learning to biological sequences---DNA, RNA and protein---has enormous potential to advance human health and environmental sustainability. To support such high-stakes applications, it is important to develop models and evaluations that not only capture underlying biology, but also have theoretical guarantees of reliability and performance. In this article, we analyze kernel methods for biological sequences, including both hand-crafted kernels and deep neural network-based kernels. We show that popular biological kernels can severely fail at learning functions or distinguishing distributions. We then develop modified kernels that (1) are universal, characteristic, and metrize the space of distributions, and (2) preserve the underlying biological inductive biases and domain knowledge embedded in the original kernel. Our results rest on novel proof techniques for kernels that handle the structure of biological sequence space--discrete, variable length sequences--and biological notions of sequence similarity. We illustrate our theoretical results in simulation and on real biological data sets.
</description>
</item>

<item>
<title>
Unified Discrete Diffusion for Categorical Data
</title>
<link>
http://jmlr.org/papers/v26/25-0171.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-0171/25-0171.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Lingxiao Zhao, Xueying Ding, Lijun Yu, Leman Akoglu</author>
<description>
Discrete diffusion models have attracted significant attention for their application to naturally discrete data, such as language and graphs. While discrete-time discrete diffusion has been established for some time, it was only recently that Campbell et al. (2022) introduced the first framework for continuous-time discrete diffusion. However, their training and backward sampling processes significantly differ from those of the discrete-time version, requiring nontrivial approximations for tractability. In this paper, we first introduce a series of generalizations and simplifications of the evidence lower bound (ELBO) that facilitate more accurate and easier optimization both discrete- and continuous-time discrete diffusion. We further establish a unification of discrete- and continuous-time discrete diffusion through shared forward process and backward parameterization. Thanks to this unification, the continuous-time diffusion can now utilize the exact and efficient backward process developed for the discrete-time case, avoiding the need for costly and inexact approximations. Similarly, the discrete-time diffusion now also employ the MCMC corrector, which was previously exclusive to the continuous-time case. Extensive experiments and ablations demonstrate the significant improvement, and we open-source our code at: https://github.com/LingxiaoShawn/USD3.
</description>
</item>

<item>
<title>
Reinforcement Learning for Infinite-Dimensional Systems
</title>
<link>
http://jmlr.org/papers/v26/24-1575.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1575/24-1575.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Wei Zhang, Jr-Shin Li</author>
<description>
Interest in reinforcement learning (RL) for large-scale systems, comprising extensive populations of intelligent agents interacting with heterogeneous environments, has surged significantly across diverse scientific domains in recent years. However, the large-scale nature of these systems often leads to high computational costs or reduced performance for most state-of-the-art RL techniques. To address these challenges, we propose a novel RL architecture and derive effective algorithms to learn optimal policies for arbitrarily large systems of agents. In our formulation, we model such systems as parameterized control systems defined on an infinite-dimensional function space. We then develop a moment kernel transform that maps the parameterized system and the value function into a reproducing kernel Hilbert space. This transformation generates a sequence of finite-dimensional moment representations for the RL problem, organized into a filtrated structure. Leveraging this RL filtration, we develop a hierarchical algorithm for learning optimal policies for the infinite-dimensional parameterized system. To enhance the algorithm&#39;s efficiency, we incorporate early stopping at each hierarchy, demonstrating the fast convergence property of the algorithm through the construction of a convergent spectral sequence. The performance and efficiency of the proposed algorithm are validated using practical examples in engineering and quantum systems.
</description>
</item>

<item>
<title>
Deep Neural Networks are Adaptive to Function Regularity and Data Distribution in Approximation and Estimation
</title>
<link>
http://jmlr.org/papers/v26/24-1148.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1148/24-1148.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Hao Liu, Jiahui Cheng, Wenjing Liao</author>
<description>
Deep learning has exhibited remarkable results across diverse areas. To understand its success, substantial research has been directed towards its theoretical foundations. Nevertheless, the majority of these studies examine how well deep neural networks can model functions with uniform regularities. In this paper, we explore a different angle: how deep neural networks can adapt to varying degrees of smoothness in functions and nonuniform data distributions across different locations and scales. More precisely, we focus on a broad class of functions defined by nonlinear tree-based approximation methods. This class encompasses a range of function types, such as functions with uniform regularities and discontinuous functions. We develop nonparametric  approximation and estimation theories  for this class using deep ReLU networks. Our results show that deep neural networks are adaptive to the nonuniform smoothness of functions and nonuniform data distributions at different locations and scales. We apply our results to several function classes, and derive the corresponding approximation and generalization errors. The validity of our results is demonstrated through  numerical experiments.
</description>
</item>

<item>
<title>
Generation of Geodesics with Actor-Critic Reinforcement Learning to Predict Midpoints
</title>
<link>
http://jmlr.org/papers/v26/24-1020.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1020/24-1020.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Kazumi Kasaura</author>
<description>
To find the shortest paths for all pairs on manifolds with infinitesimally defined metrics, we introduce a framework to generate them by predicting midpoints recursively. To learn midpoint prediction, we propose an actor-critic approach. We prove the soundness of our approach and show experimentally that the proposed method outperforms existing methods on several planning tasks, including path planning for agents with complex kinematics and motion planning for multi-degree-of-freedom robot arms.
</description>
</item>

<item>
<title>
Learning-to-Optimize with PAC-Bayesian Guarantees: Theoretical Considerations and Practical Implementation
</title>
<link>
http://jmlr.org/papers/v26/24-0486.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0486/24-0486.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Michael Sucker, Jalal Fadili, Peter Ochs</author>
<description>
We use the PAC-Bayesian theory for the setting of learning-to-optimize. To the best of our knowledge, we present the first framework to learn optimization algorithms with provable generalization guarantees (PAC-Bayesian bounds) and explicit trade-off between convergence guarantees and convergence speed, which contrasts with the typical worst-case analysis. Our learned optimization algorithms provably outperform related ones derived from a worst-case analysis. The results rely on PAC-Bayesian bounds for general, possibly unbounded loss-functions based on exponential families. Further, we provide a concrete algorithmic realization of the framework and new methodologies for learning-to-optimize. Finally, we conduct four practically relevant experiments to support our theory. With this, we showcase that the provided learning framework yields optimization algorithms that provably outperform the state-of-the-art by orders of magnitude.
</description>
</item>

<item>
<title>
Sparse Semiparametric Discriminant Analysis for High-dimensional Zero-inflated Data
</title>
<link>
http://jmlr.org/papers/v26/24-0046.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0046/24-0046.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Hee Cheol Chung, Yang Ni, Irina Gaynanova</author>
<description>
Sequencing-based technologies provide an abundance of high-dimensional biological data sets with highly skewed and zero-inflated measurements. Despite the computational efficiency and high interpretability offered by linear classification methods, the violation of underlying distribution assumptions, driven by high skewness and zero inflation, results in invalid classification rules and interpretations. Furthermore, existing data transformation methods addressing these violations introduce ambiguity, rendering the final model and classification performance contingent on the specific transformation employed. To tackle these challenges, we propose a novel semiparametric framework for discriminant analysis based on the truncated latent Gaussian copula model. This model accommodates skewness and zero inflation, and its estimation procedure ensures robustness against data transformations. To facilitate model interpretability, we incorporate $\ell_1$ sparsity regularization and establish the consistency of the classification directions in high-dimensional settings. We validate our approach using human gut microbiome, breast cancer microRNA, and single-cell RNA sequencing data, highlighting its superior classification accuracy and robustness to data transformations.
</description>
</item>

<item>
<title>
Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
</title>
<link>
http://jmlr.org/papers/v26/23-1605.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1605/23-1605.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Michael Albergo, Nicholas M. Boffi, Eric Vanden-Eijnden</author>
<description>
A class of generative models that unifies flow-based and diffusion-based methods is introduced. These models extend the framework proposed in Albergo and Vanden-Eijnden (2023), enabling the use of a broad class of continuous-time stochastic processes called stochastic interpolants to bridge any two probability density functions exactly in finite time. These interpolants are built by combining data from the two prescribed densities with an additional latent variable that shapes the bridge in a flexible way. The time-dependent density function of the interpolant is shown to satisfy a transport equation as well as a family of forward and backward Fokker-Planck equations with tunable diffusion coefficient. Upon consideration of the time evolution of an individual sample, this viewpoint leads to both deterministic and stochastic generative models based on probability flow equations or stochastic differential equations with an adjustable level of noise. The drift coefficients entering these models are time-dependent velocity fields characterized as the unique minimizers of simple quadratic objective functions, one of which is a new objective for the score. We show that minimization of these quadratic objectives leads to control of the likelihood for generative models built upon stochastic dynamics, while likelihood control for deterministic dynamics is more stringent. We also construct estimators for the likelihood and the cross entropy of interpolant-based generative models, and we discuss connections with other methods such as score-based diffusion models, stochastic localization, probabilistic denoising, and rectifying flows. In addition, we demonstrate that stochastic interpolants recover the Schrödinger bridge between the two target densities when explicitly optimizing over the interpolant. Finally, algorithmic aspects are discussed and the approach is illustrated on numerical examples.
</description>
</item>

<item>
<title>
Efficient Methods for Non-stationary Online Learning
</title>
<link>
http://jmlr.org/papers/v26/23-1188.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1188/23-1188.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Peng Zhao, Yan-Feng Xie, Lijun Zhang, Zhi-Hua Zhou</author>
<description>
Non-stationary online learning has drawn much attention in recent years. In particular, dynamic regret and adaptive regret are proposed as two principled performance measures for online convex optimization in non-stationary environments. To optimize them, a two-layer online ensemble is usually deployed due to the inherent uncertainty of non-stationarity, in which multiple base-learners are maintained and a meta-algorithm is employed to track the best one on the fly. However, the two-layer structure raises concerns about computational complexity --- such methods typically maintain $O(\log T)$ base-learners simultaneously for a $T$-round online game and thus perform multiple projections onto the feasible domain per round, which becomes the computational bottleneck when the domain is complicated. In this paper, we present efficient methods for optimizing dynamic regret and adaptive regret that reduce the number of projections per round from $O(\log T)$ to $1$. The proposed algorithms require only one gradient query and one function evaluation at each round. Our technique hinges on the reduction mechanism developed in parameter-free online learning and requires non-trivial modifications for non-stationary online methods. Furthermore, we study an even stronger measure, namely &#34;interval dynamic regret&#34;, and reduce the number of projections per round from $O(\log^2 T)$ to $1$ for minimizing it. Our reduction demonstrates broad generality and applies to two important applications: online stochastic control and online principal component analysis, resulting in methods that are both efficient and optimal. Finally, empirical studies verify our theoretical findings.
</description>
</item>

<item>
<title>
Decentralized Asynchronous Optimization with DADAO allows Decoupling and Acceleration
</title>
<link>
http://jmlr.org/papers/v26/23-0748.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0748/23-0748.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Adel Nabli, Edouard Oyallon</author>
<description>
DADAO is the first decentralized, accelerated, asynchronous, primal, first-order algorithm to minimize  a sum of $L$-smooth and $\mu$-strongly convex functions distributed over a network of size $n$.  Modeling the gradient updates and gossip communication procedures with separate independent Poisson Point Processes allows us to decouple the computation and communication steps, which can be run in parallel, while making the whole approach completely asynchronous. This leads to communication acceleration compared to synchronous approaches. Our method employs primal gradients and avoids using a multi-consensus inner loop and other ad-hoc mechanisms. By relating the smallest positive eigenvalue $1/\chi_1$ of the Laplacian matrix $\Lambda$  and the maximal resistance $\chi_2\leq \chi_1$ of the graph to a sufficient minimal communication rate, we show that  DADAO requires $\mathcal{O}(n\sqrt{\frac{L}{\mu}}\log(\frac{1}{\epsilon}))$ local gradients and only $\mathcal{O}(\sqrt{\chi_1\chi_2}\operatorname{Tr}\Lambda\sqrt{\frac{L}{\mu}}\log(\frac{1}{\epsilon}))$ communications to reach $\epsilon$-precision, up to logarithmic terms. Thus, we simultaneously obtain an accelerated rate for computations and communications, leading to an improvement over state-of-the-art works, our simulations further validating the strength of our relatively unconstrained method. Moreover, we propose a SDP relaxation to find the gossip rate of each edge minimizing the total number of communications for a given graph, resulting in faster convergence compared to standard approaches relying on uniform communication weights.
</description>
</item>

<item>
<title>
Mixtures of Gaussian Process Experts with SMC^2
</title>
<link>
http://jmlr.org/papers/v26/22-0973.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0973/22-0973.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Teemu Härkönen, Sara Wade, Kody Law, Lassi Roininen</author>
<description>
Gaussian processes are a key component of many flexible statistical and machine learning models. However, they exhibit cubic computational complexity and high memory constraints due to the need of inverting and storing a full covariance matrix. To circumvent this, mixtures of Gaussian process experts have been considered where data points are assigned to independent experts, reducing the complexity by allowing inference based on smaller, local covariance matrices. Moreover, mixtures of Gaussian process experts substantially enrich the model&#39;s flexibility, allowing for behaviors such as non-stationarity, heteroscedasticity, and discontinuities. In this work, we construct a novel inference approach based on nested sequential Monte Carlo samplers to simultaneously infer both the gating network and Gaussian process expert parameters. This greatly improves inference  compared to importance sampling, particularly in settings when a stationary Gaussian process is inappropriate,  while still being thoroughly parallelizable.
</description>
</item>

<item>
<title>
Robust Point Matching with Distance Profiles
</title>
<link>
http://jmlr.org/papers/v26/24-2224.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-2224/24-2224.pdf
</pdf>
<pubDate>2025</pubDate>
<author>YoonHaeng Hur, Yuehaw Khoo</author>
<description>
Computational difficulty of quadratic matching and the Gromov-Wasserstein distance has led to various approximation and relaxation schemes. One of such methods, relying on the notion of distance profiles, has been widely used in practice, but its theoretical understanding is limited. By delving into the statistical complexity of the previously proposed method based on distance profiles, we show that it suffers from the curse of dimensionality unless we make certain assumptions on the underlying metric measure spaces. Building on this insight, we propose and analyze a modified matching procedure that can be used to robustly match points under a certain probabilistic setting. We demonstrate the performance of the proposed methods using simulations and real data applications to complement the theoretical findings. As a result, we contribute to the literature by providing theoretical underpinnings of the matching procedures based on distance invariants like distance profiles, which have been widely used in practice but rarely analyzed theoretically.
</description>
</item>

<item>
<title>
BoFire: Bayesian Optimization Framework Intended for Real Experiments
</title>
<link>
http://jmlr.org/papers/v26/24-1540.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1540/24-1540.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Johannes P. Dürholt, Thomas S. Asche, Johanna Kleinekorte, Gabriel Mancino-Ball, Benjamin Schiller, Simon Sung, Julian Keupp, Aaron Osburg, Toby Boyne, Ruth Misener, Rosona Eldred, Chrysoula Kappatou, Robert M. Lee, Dominik Linzner, Wagner Steuer Costa, David Walz, Niklas Wulkow, Behrang Shafei</author>
<description>
Our open-source Python package BoFire combines Bayesian Optimization (BO) with other design of experiments (DoE) strategies focusing on developing and optimizing new chemistry. Previous BO implementations, for example as they exist in the literature or software, require substantial adaptation for effective real-world deployment in chemical industry. BoFire provides a rich feature-set with extensive configurability and realizes our vision of fast-tracking research contributions into industrial use via maintainable open-source software. Owing to quality-of-life features like JSON-serializability of problem formulations, BoFire enables seamless integration of BO into RESTful APIs, a common architecture component for both self-driving laboratories and human-in-the-loop setups. This paper discusses the differences between BoFire and other BO implementations and outlines ways that BO research needs to be adapted for real-world use in a chemistry setting.
</description>
</item>

<item>
<title>
Reliever: Relieving the Burden of Costly Model Fits for Changepoint Detection
</title>
<link>
http://jmlr.org/papers/v26/24-1108.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1108/24-1108.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Chengde Qian, Guanghui Wang, Changliang Zou</author>
<description>
Changepoint detection typically relies on a grid-search strategy for optimal data segmentation. When model fitting itself is expensive, repeatedly fitting a model on every candidate segment dominates the computation. Existing approaches mitigate this by pruning the grid, thus reducing the number of segments (and model fits). We propose Reliever, which instead cuts the number of model fits directly and nests seamlessly within standard grid-search routines. Reliever fits a small, deterministic collection of proxy models and reuses them wherever they apply, making it compatible with a wide range of existing algorithms. For high-dimensional regression with changepoints, coupling Reliever with an optimal grid-search method yields changepoint and coefficient estimators that are rate-optimal up to a logarithmic factor. Extensive numerical experiments demonstrate that Reliever rapidly and accurately detects changepoints across a wide range of high-dimensional and nonparametric models.
</description>
</item>

<item>
<title>
Variational Inference for Uncertainty Quantification: an Analysis of Trade-offs
</title>
<link>
http://jmlr.org/papers/v26/24-0878.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0878/24-0878.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Charles C. Margossian, Loucas Pillaud-Vivien, Lawrence K. Saul</author>
<description>
Given an intractable distribution $p$, the problem of variational inference (VI) is to find the best approximation from some more tractable family $Q$. Commonly, one chooses $Q$ to be a family of factorized distributions (i.e., the mean-field assumption), even though $p$ itself does not factorize. We show that this mismatch leads to an impossibility theorem: if $p$ does not factorize, then any factorized approximation $q\!\in\!Q$ can correctly estimate at most one of the following three measures of uncertainty: (i) the marginal variances, (ii) the marginal precisions, or (iii) the generalized variance (which for elliptical distributions is closely related to the entropy). In practice, the best variational approximation in $Q$ is found by minimizing some divergence $D(q,p)$ between distributions, and so we ask: how does the choice of divergence determine which measure of uncertainty, if any, is correctly estimated by VI? We consider the classic Kullback-Leibler divergences, the more general $\alpha$-divergences, and a score-based divergence which compares $\nabla \log p$ and $\nabla \log q$. We thoroughly analyze the case where $p$ is a Gaussian and $q$ is a (factorized) Gaussian. In this setting, we show that all the considered divergences can be ordered based on the estimates of uncertainty they yield as objective functions for VI. Finally, we empirically evaluate the validity of this ordering when the target distribution $p$ is not Gaussian.
</description>
</item>

<item>
<title>
Are Ensembles Getting Better All the Time?
</title>
<link>
http://jmlr.org/papers/v26/24-0408.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0408/24-0408.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Pierre-Alexandre Mattei, Damien Garreau</author>
<description>
Ensemble methods combine the predictions of several base models. We study whether or not including more models always improves their average performance. This question depends on the kind of ensemble considered, as well as the predictive metric chosen. We focus on situations where all members of the ensemble are a priori expected to perform equally well, which is the case of several popular methods such as random forests or deep ensembles. In this setting, we show that ensembles are getting better all the time if, and only if, the considered loss function is convex. More precisely, in that case, the loss of the ensemble is a decreasing function of the number of models. When the loss function is nonconvex, we show a series of results that can be summarised as: ensembles of good models keep getting better, and ensembles of bad models keep getting worse. To this end, we prove a new result on the monotonicity of tail probabilities that may be of independent interest. We illustrate our results on a  medical problem (diagnosing melanomas using neural nets) and a “wisdom of crowds” experiment (guessing the ratings of upcoming movies).
</description>
</item>

<item>
<title>
An Adaptive Parameter-free and Projection-free Restarting Level Set Method for Constrained Convex Optimization Under the Error Bound Condition
</title>
<link>
http://jmlr.org/papers/v26/24-0201.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0201/24-0201.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Qihang Lin, Negar Soheili, Runchao Ma, Selvaprabu Nadarajah</author>
<description>
Recent efforts to accelerate first-order methods have focused on convex optimization problems that satisfy a geometric property known as error-bound condition, which covers a broad class of problems, including piece-wise linear programs and strongly convex programs. Parameter-free first-order methods that employ projection-free updates have the potential to broaden the benefit of acceleration. Such a method has been developed for unconstrained convex optimization but is lacking for general constrained convex optimization. We propose a parameter-free level-set method for the latter constrained case based on projection-free subgradient method that exhibits accelerated convergence for problems that satisfy an error-bound condition. Our method maintains a separate copy of the level-set sub-problem for each level parameter value and restarts the computation of these copies based on objective function progress. Applying such a restarting scheme in a level-set context is novel and results in an algorithm that dynamically adapts the precision of each copy. This property is key to extending prior restarting methods based on static precision that have been proposed for unconstrained convex optimization to handle constraints. We report promising numerical performance relative to benchmark methods.
</description>
</item>

<item>
<title>
Operator Learning for Hyperbolic PDEs
</title>
<link>
http://jmlr.org/papers/v26/23-1724.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1724/23-1724.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Christopher Wang, Alex Townsend</author>
<description>
We construct the first rigorously justified probabilistic algorithm for recovering the solution operator of a hyperbolic partial differential equation (PDE) in two variables from input-output training pairs. The primary challenge of recovering the solution operator of hyperbolic PDEs is the presence of characteristics, along which the associated Green&#39;s function is discontinuous. Therefore, a central component of our algorithm is a rank detection scheme that identifies the approximate location of the characteristics. By combining the randomized singular value decomposition with an adaptive hierarchical partition of the domain, we construct an approximant to the solution operator using $O(\Psi_\epsilon^{-1}\epsilon^{-7}\log(\Xi_\epsilon^{-1}\epsilon^{-1}))$ input-output pairs with relative error $O(\Xi_\epsilon^{-1}\epsilon)$ in the operator norm as $\epsilon\to0$, with high probability. Here, $\Psi_\epsilon$ represents the existence of degenerate singular values of the solution operator, and $\Xi_\epsilon$ measures the quality of the training data. Our assumptions on the regularity of the coefficients of the hyperbolic PDE are relatively weak given that hyperbolic PDEs do not have the &#34;instantaneous smoothing effect&#34; of elliptic and parabolic PDEs, and our recovery rate improves as the regularity of the coefficients increases. We also include numerical experiments which corroborate our theoretical findings.
</description>
</item>

<item>
<title>
Optimal subsampling for high-dimensional partially linear models via machine learning methods
</title>
<link>
http://jmlr.org/papers/v26/23-1475.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1475/23-1475.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yujing Shao, Lei Wang, Heng Lian, Haiying Wang</author>
<description>
In this paper, we explore optimal subsampling strategies for estimating the parametric regression coefficients in partially linear models with unknown nuisance functions involving high-dimensional and potentially endogenous covariates. To address model misspecifications and the curse of dimensionality, we leverage flexible machine learning (ML) techniques to estimate the unknown nuisance functions.  By constructing an unbiased subsampling Neyman-orthogonal score function, we eliminate regularization bias. A two-step algorithm is then used to obtain appropriate ML estimators of the nuisance functions, mitigating the risk of over-fitting. Using martingale techniques, we establish the unconditional consistency and asymptotic normality of the subsample estimators. Furthermore, we derive optimal subsampling probabilities, including A-optimal and L-optimal probabilities as special cases. The proposed optimal subsampling approach is extended to partially linear instrumental variable models to account for potential endogeneity through instrumental variables.  Simulation studies and an empirical analysis of the Physicochemical Properties of Protein Tertiary Structure dataset demonstrate the superior performance of our subsample estimators.
</description>
</item>

<item>
<title>
Decentralized Sparse Linear Regression via Gradient-Tracking
</title>
<link>
http://jmlr.org/papers/v26/23-1168.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1168/23-1168.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Marie Maros, Gesualdo Scutari, Ying Sun, Guang Cheng</author>
<description>
We study   sparse linear regression over a network of agents, modeled as an undirected  graph without a center node. The estimation of the $s$-sparse parameter is  formulated as a constrained LASSO problem wherein each agent   owns a subset of the $N$ total   observations. We analyze the convergence rate and statistical guarantees of a distributed projected gradient tracking-based algorithm under high-dimensional scaling, allowing the ambient  dimension $d$ to grow with (and possibly exceed) the sample size $N$. Our theory shows that, under standard   notions of restricted strong convexity and smoothness of the average loss functions,  suitable conditions on the network connectivity  and algorithm tuning, the distributed algorithm converges   globally at a   linear rate to an estimate that is within the centralized   statistical precision  of the model, $O(s\log d/N)$. When $s\log d/N=o(1)$, a condition necessary for statistical consistency, an $\varepsilon$-optimal solution is attained after  ${O}(\kappa \log (1/\varepsilon))$ gradient computations  and $O(\kappa/(1-\rho) \log (1/\varepsilon))$  communication rounds,
where $\kappa$ is the restricted condition number of the loss function and $\rho$ measures the network connectivity. The computation cost matches that of  the centralized projected gradient algorithm despite  having data distributed; whereas the communication rounds reduce as the network connectivity improves. Overall, our study   reveals  interesting connections between statistical efficiency, network connectivity and topology, and  convergence rate in  the high dimensional setting.
</description>
</item>

<item>
<title>
Calibrated Inference: Statistical Inference that Accounts for Both Sampling Uncertainty and Distributional Uncertainty
</title>
<link>
http://jmlr.org/papers/v26/23-0714.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0714/23-0714.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yujin Jeong, Dominik Rothenhäusler</author>
<description>
How can we draw trustworthy scientific conclusions? One criterion is that a study can be replicated by independent teams. While replication is critically important, it is arguably insufficient. If a study is biased for some reason and other studies recapitulate the approach then findings might be consistently incorrect. It has been argued that trustworthy scientific conclusions require disparate sources of evidence. However, different methods might have shared biases, making it difficult to judge the trustworthiness of a result. We formalize this issue by introducing a &#34;distributional uncertainty model&#34;, wherein dense distributional shifts emerge as the superposition of numerous small random changes. The distributional perturbation model arises under a symmetry assumption on distributional shifts and is strictly weaker than assuming that the data is i.i.d. from the target distribution. We show that a stability analysis on a single data set allows us to construct confidence intervals that account for both sampling uncertainty and distributional uncertainty.
</description>
</item>

<item>
<title>
Relaxed Gaussian Process Interpolation: a Goal-Oriented Approach to Bayesian Optimization
</title>
<link>
http://jmlr.org/papers/v26/22-0828.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0828/22-0828.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Sébastien J. Petit, Julien Bect, Emmanuel Vazquez</author>
<description>
This work presents a new procedure for obtaining predictive distributions in the context of Gaussian process (GP) modeling, with a relaxation of the interpolation constraints outside ranges of interest: the mean of the predictive distribution no longer necessarily interpolates the observed values when they are outside ranges of interest, but is simply constrained to remain outside. This method called relaxed Gaussian process (reGP) interpolation provides better predictive distributions in ranges of interest, especially in cases where a stationarity assumption for the GP model is not appropriate. It can be viewed as a goal-oriented method and becomes particularly interesting in Bayesian optimization, for example, for the minimization of an objective function, where good predictive distributions for low function values are important. When the expected improvement criterion and reGP are used for sequentially choosing evaluation points, the convergence of the resulting optimization algorithm is theoretically guaranteed (provided that the function to be optimized lies in the reproducing kernel Hilbert space attached to the known covariance of the underlying Gaussian process). Experiments indicate that using reGP instead of stationary GP models in Bayesian optimization is beneficial.
</description>
</item>

<item>
<title>
Linear Separation Capacity of Self-Supervised Representation Learning
</title>
<link>
http://jmlr.org/papers/v26/24-2032.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-2032/24-2032.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Shulei Wang</author>
<description>
Recent advances in self-supervised learning have highlighted the efficacy of data augmentation in learning data representation from unlabeled data. Training a linear model atop these enhanced representations can yield an adept classifier. Despite the remarkable empirical performance, the underlying mechanisms that enable data augmentation to unravel nonlinear data structures into linearly separable representations remain elusive. This paper seeks to bridge this gap by investigating under what conditions learned representations can linearly separate manifolds when data is drawn from a multi-manifold model. Our investigation reveals that data augmentation offers additional information beyond observed data and can thus improve the information-theoretic optimal rate of linear separation capacity. In particular, we show that self-supervised learning can linearly separate manifolds with a smaller distance than unsupervised learning, underscoring the additional benefits of data augmentation. Our theoretical analysis further underscores that the performance of downstream linear classifiers primarily hinges on the linear separability of data representations rather than the size of the labeled data set, reaffirming the viability of constructing efficient classifiers with limited labeled data amid an expansive unlabeled data set.
</description>
</item>

<item>
<title>
On the Convergence of Projected Policy Gradient for Any Constant Step Sizes
</title>
<link>
http://jmlr.org/papers/v26/24-1530.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1530/24-1530.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jiacai Liu, Wenye Li, Dachao Lin, Ke Wei, Zhihua Zhang</author>
<description>
Projected policy gradient (PPG) is a basic policy optimization method in reinforcement learning. Given access to exact policy  evaluations, previous studies have established the sublinear convergence of PPG for sufficiently small step sizes based on the smoothness and the gradient domination properties of the value function. However, as the step size  goes to infinity, PPG reduces to the classic policy iteration method, which suggests the convergence of PPG even for large step sizes. In this paper, we fill this gap and show that PPG admits a sublinear convergence for any constant step sizes. Due to the existence of the state-wise visitation measure in the expression of policy gradient, the existing optimization-based analysis framework for a preconditioned version of PPG (i.e., projected Q-ascent) is not applicable, to the best of our knowledge. Instead, we proceed the proof by computing the state-wise improvement lower bound of PPG based on its inherent structure. In addition, the finite iteration convergence of PPG for any constant step size is further established, which is also new.
</description>
</item>

<item>
<title>
Learning with Linear Function Approximations in Mean-Field Control
</title>
<link>
http://jmlr.org/papers/v26/24-1221.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1221/24-1221.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Erhan Bayraktar, Ali Devran Kara</author>
<description>
The paper focuses on mean-field type multi-agent control problems with finite state and action spaces where the dynamics and cost structures are symmetric and homogeneous, and are affected by the distribution of the agents. A standard solution method for these problems is to consider the infinite population limit as an approximation and use symmetric solutions of the limit problem to achieve near optimality. The control policies, and in particular the dynamics, depend on the population distribution in the finite population setting, or the marginal distribution of the state variable of a representative agent for the infinite population setting. Hence, learning and planning for these control problems generally require estimating the reaction of the system to all possible state distributions of the agents. To overcome this issue, we consider linear function approximation for the control problem and provide coordinated and independent learning methods. We rigorously establish error upper bounds for the performance of learned solutions. The performance gap stems from (i) the mismatch due to estimating the true model with a linear one, and (ii) using the infinite population solution in the finite population problem as an approximate control. The provided upper bounds quantify the impact of these error sources on the overall performance.
</description>
</item>

<item>
<title>
A New Random Reshuffling Method for Nonsmooth Nonconvex Finite-sum Optimization
</title>
<link>
http://jmlr.org/papers/v26/24-0891.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0891/24-0891.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Junwen Qiu, Xiao Li, Andre Milzarek</author>
<description>
Random reshuffling techniques are prevalent in large-scale applications, such as training neural networks. While the convergence and acceleration effects of random reshuffling-type methods are fairly well understood in the smooth setting, much less studies seem available in the nonsmooth case. In this work, we design a new normal map-based proximal random reshuffling (norm-PRR) method for nonsmooth nonconvex finite-sum problems. We show that norm-PRR achieves the iteration complexity ${\cal O}(n^{-1/3}T^{-2/3})$ where $n$ denotes the number of component functions $f(\cdot,i)$ and $T$ counts the total number of iterations. This improves the currently known complexity bounds for this class of problems by a factor of $n^{-1/3}$ in terms of the number of gradient evaluations.  Additionally, we prove that norm-PRR converges linearly under the (global) Polyak-Łojasiewicz condition and in the interpolation setting. We further complement these non-asymptotic results and provide an in-depth analysis of the asymptotic properties of norm-PRR. Specifically, under the (local) Kurdyka-Łojasiewicz  inequality, the whole sequence of iterates generated by norm-PRR is shown to converge to a single stationary point. Moreover, we derive last-iterate convergence rates that can match those in the smooth, strongly convex setting. Finally, numerical experiments are performed on nonconvex classification tasks to illustrate the efficiency of the proposed approach.
</description>
</item>

<item>
<title>
Model-free Change-Point Detection Using AUC of a Classifier
</title>
<link>
http://jmlr.org/papers/v26/24-0365.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0365/24-0365.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Rohit Kanrar, Feiyu Jiang, Zhanrui Cai</author>
<description>
In contemporary data analysis, it is increasingly common to work with non-stationary complex data sets. These data sets typically extend beyond the classical low-dimensional Euclidean space, making it challenging to detect shifts in their distribution without relying on strong structural assumptions. This paper proposes a novel offline change-point detection method that leverages classifiers developed in the statistics and machine learning community.  With suitable data splitting, the test statistic is constructed through sequential computation of the Area Under the Curve (AUC) of a classifier, which is trained on data segments on both ends of the sequence. It is shown that the resulting AUC process attains its maxima at the true change-point location, which facilitates the change-point estimation. The proposed method is characterized by its complete nonparametric nature, high versatility, considerable flexibility, and absence of stringent assumptions on the underlying data or any distributional shifts. Theoretically, we derive the limiting pivotal distribution of the proposed test statistic under null, as well as the asymptotic behaviors under both local and fixed alternatives. The localization rate of the change-point estimator is also provided. Extensive simulation studies and the analysis of two real-world data sets illustrate the superior performance of our approach compared to existing model-free change-point detection methods.
</description>
</item>

<item>
<title>
EF21 with Bells &amp; Whistles: Six Algorithmic Extensions of Modern Error Feedback
</title>
<link>
http://jmlr.org/papers/v26/24-0059.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0059/24-0059.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Ilyas Fatkhullin, Igor Sokolov, Eduard Gorbunov, Zhize Li, Peter Richtárik</author>
<description>
First proposed by Seide (2014) as a heuristic, error feedback (EF) is a very popular mechanism for enforcing convergence of distributed gradient-based optimization methods enhanced with communication compression strategies based on the application of contractive compression operators. However, existing theory of EF relies on very strong assumptions (e.g., bounded gradients), and provides pessimistic convergence rates (e.g., while the best known rate for EF in the smooth nonconvex regime, and when full gradients are compressed, is $O(1/T^{2/3})$, the rate of gradient descent in the same regime is $O(1/T)$). Recently, Richtàrik et al. (2021) proposed a new error feedback mechanism, EF21, based on the construction of a Markov compressor induced by a contractive compressor. EF21 removes the aforementioned theoretical deficiencies of EF and at the same time works better in practice. In this work we propose six practical extensions of EF21, all supported by strong convergence theory: partial participation, stochastic approximation, variance reduction, proximal setting, momentum, and bidirectional compression. To the best of our knowledge, several of these techniques have not been previously analyzed in combination with EF, and in cases where prior analysis exists---such as for bidirectional compression---our theoretical convergence guarantees significantly improve upon existing results.
</description>
</item>

<item>
<title>
Multiple Instance Verification
</title>
<link>
http://jmlr.org/papers/v26/23-1590.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1590/23-1590.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Xin Xu, Eibe Frank, Geoffrey Holmes</author>
<description>
We explore multiple instance verification, a problem setting in which a query instance is verified against a bag of target instances with heterogeneous, unknown relevancy. We show that naive adaptations of attention-based multiple instance learning (MIL) methods and standard verification methods like Siamese neural networks are unsuitable for this setting: directly combining state-of-the-art (SOTA) MIL methods and Siamese networks is shown to be no better, and sometimes significantly worse, than a simple baseline model. Postulating that this may be caused by the failure of the representation of the target bag to incorporate the query instance, we introduce a new pooling approach named “cross-attention pooling” (CAP). Under the CAP framework, we propose two novel attention functions to address the challenge of distinguishing between highly similar instances in a target bag. Through empirical studies on three different verification tasks, we demonstrate that CAP outperforms adaptations of SOTA MIL methods and the baseline by substantial margins, in terms of both classification accuracy and the ability to detect key instances. The superior ability to identify key instances is attributed to the new attention functions by ablation studies.
</description>
</item>

<item>
<title>
Learning from Similar Linear Representations: Adaptivity, Minimaxity, and Robustness
</title>
<link>
http://jmlr.org/papers/v26/23-0902.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0902/23-0902.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Ye Tian, Yuqi Gu, Yang Feng</author>
<description>
Representation multi-task learning (MTL) has achieved tremendous success in practice. However, the theoretical understanding of these methods is still lacking. Most existing theoretical works focus on cases where all tasks share the same representation, and claim that MTL almost always improves performance. Nevertheless, as the number of tasks grows, assuming all tasks share the same representation is unrealistic. Furthermore, empirical findings often indicate that a shared representation does not necessarily improve single-task learning performance. In this paper, we aim to understand how to learn from tasks with similar but not exactly the same linear representations, while dealing with outlier tasks. Assuming a known intrinsic dimension, we propose a penalized empirical risk minimization method and a spectral method that are adaptive to the similarity structure and robust to outlier tasks. Both algorithms outperform single-task learning when representations across tasks are sufficiently similar and the proportion of outlier tasks is small. Moreover, they always perform at least as well as single-task learning, even when the representations are dissimilar. We provide information-theoretic lower bounds to demonstrate that both methods are nearly minimax optimal in a large regime, with the spectral method being optimal in the absence of outlier tasks. Additionally, we introduce a thresholding algorithm to adapt to an unknown intrinsic dimension. We conduct extensive numerical experiments to validate our theoretical findings.
</description>
</item>

<item>
<title>
Exponential Family Graphical Models: Correlated Replicates and Unmeasured Confounders, with Applications to fMRI Data
</title>
<link>
http://jmlr.org/papers/v26/22-1421.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-1421/22-1421.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yanxin Jin, Yang Ning, Kean Ming Tan</author>
<description>
Graphical models have been used extensively for modeling brain connectivity networks. However, unmeasured confounders and correlations among measurements are often overlooked during model fitting, which may lead to spurious scientific discoveries. Motivated by functional magnetic resonance imaging (fMRI) studies, we propose a novel method for constructing brain connectivity networks with correlated replicates and latent effects. In a typical fMRI study, each participant is scanned and fMRI measurements are collected across a period of time. In many cases, subjects may have different states of mind that cannot be measured during the brain scan: for instance, some subjects may be awake during the first half of the brain scan, and may fall asleep during the second half of the brain scan. To model the correlation among replicates and latent effects induced by the different states of mind, we assume that the correlated replicates within each independent subject follow a one-lag vector autoregressive model, and that the latent effects induced by the unmeasured confounders are piecewise constant. Theoretical guarantees are established for parameter estimation. We demonstrate via extensive numerical studies that our method is able to estimate latent variable graphical models with correlated replicates more accurately than existing methods.
</description>
</item>

<item>
<title>
Optimizing Return Distributions with Distributional Dynamic Programming
</title>
<link>
http://jmlr.org/papers/v26/25-0210.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-0210/25-0210.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Bernardo Ávila Pires, Mark Rowland, Diana Borsa, Zhaohan Daniel Guo, Khimya Khetarpal, André Barreto, David Abel, Rémi Munos, Will Dabney</author>
<description>
We introduce distributional dynamic programming (DP) methods for optimizing statistical functionals of the return distribution, with standard reinforcement learning as a special case. Previous distributional DP methods could optimize the same class of expected utilities as classic DP. To go beyond, we combine distributional DP with stock augmentation, a technique previously introduced for classic DP in the context of risk-sensitive RL, where the MDP state is augmented with a statistic of the rewards obtained since the first time step. We find that a number of recently studied problems can be formulated as stock-augmented return distribution optimization, and we show that we can use distributional DP to solve them. We analyze distributional value and policy iteration, with bounds and a study of what objectives these distributional DP methods can or cannot optimize. We describe a number of applications outlining how to use distributional DP to solve different stock-augmented return distribution optimization problems, for example maximizing conditional value-at-risk, and homeostatic regulation. To highlight the practical potential of stock-augmented return distribution optimization and distributional DP, we introduce an agent that combines DQN and the core ideas of distributional DP, and empirically evaluate it for solving instances of the applications discussed.
</description>
</item>

<item>
<title>
Imprecise Multi-Armed Bandits: Representing Irreducible Uncertainty as a Zero-Sum Game
</title>
<link>
http://jmlr.org/papers/v26/24-2001.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-2001/24-2001.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Vanessa Kosoy</author>
<description>
We introduce a novel multi-armed bandit framework, where each arm is associated with a fixed unknown credal set over the space of outcomes (which can be richer than just the reward). The arm-to-credal-set correspondence comes from a known class of hypotheses. We then define a notion of regret corresponding to the lower prevision defined by these credal sets. Equivalently, the setting can be regarded as a two-player zero-sum game, where, on each round, the agent chooses an arm and the adversary chooses the distribution over outcomes from a set of options associated with this arm. The regret is defined with respect to the value of game. For certain natural hypothesis classes, loosely analogous to stochastic linear bandits (which are a special case of the resulting setting), we propose an algorithm and prove a corresponding upper bound on regret.
</description>
</item>

<item>
<title>
Early Alignment in Two-Layer Networks Training is a Two-Edged Sword
</title>
<link>
http://jmlr.org/papers/v26/24-1523.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1523/24-1523.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Etienne Boursier, Nicolas Flammarion</author>
<description>
Training neural networks with first order optimisation methods is at the core of the empirical success of deep learning. The scale of initialisation is a crucial factor, as small initialisations are generally associated to a feature learning regime, for which gradient descent is implicitly biased towards simple solutions. This work provides a general and quantitative description of the early alignment phase, originally introduced by Maennel et al. (2018). For small initialisation and one hidden ReLU layer networks, the early stage of the training dynamics leads to an alignment of the neurons towards key directions. This alignment induces a sparse representation of the network, which is directly related to the implicit bias of gradient flow  at convergence. This sparsity inducing alignment however comes at the expense of difficulties in minimising the training objective: we also provide a simple data example for which overparameterised networks fail to converge towards global minima and only converge to a spurious stationary point instead.
</description>
</item>

<item>
<title>
Hierarchical Decision Making Based on Structural Information Principles
</title>
<link>
http://jmlr.org/papers/v26/24-1184.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1184/24-1184.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Xianghua Zeng, Hao Peng, Dingli Su, Angsheng Li</author>
<description>
Hierarchical Reinforcement Learning (HRL) is a promising approach for managing task complexity across multiple levels of abstraction and accelerating long-horizon agent exploration. However, the effectiveness of hierarchical policies heavily depends on prior knowledge and manual assumptions about skill definitions and task decomposition. In this paper, we propose a novel Structural Information principles-based framework, namely SIDM, for hierarchical Decision Making in both single-agent and multi-agent scenarios. Central to our work is the utilization of structural information embedded in the decision-making process to adaptively and dynamically discover and learn hierarchical policies through environmental abstractions. Specifically, we present an abstraction mechanism that processes historical state-action trajectories to construct abstract representations of states and actions. We define and optimize directed structural entropy—a metric quantifying the uncertainty in transition dynamics between abstract states—to discover skills that capture key transition patterns in RL environments. Building on these findings, we develop a skill-based learning method for single-agent scenarios and a role-based collaboration method for multi-agent scenarios, both of which can flexibly integrate various underlying algorithms for enhanced performance. Extensive evaluations on challenging benchmarks demonstrate that our framework significantly and consistently outperforms state-of-the-art baselines, improving the effectiveness, efficiency, and stability of policy learning by up to 32.70%, 64.86%, and 88.26%, respectively, as measured by average rewards, convergence timesteps, and standard deviations.
</description>
</item>

<item>
<title>
Generative Adversarial Networks: Dynamics
</title>
<link>
http://jmlr.org/papers/v26/24-0848.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0848/24-0848.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Matias G. Delgadino, Bruno B. Suassuna, Rene Cabrera</author>
<description>
We study quantitatively the overparametrization limit of the original Wasserstein-GAN algorithm. Effectively, we show that the algorithm is a stochastic discretization of a system of continuity equations for the parameter distributions of the generator and discriminator. We show that parameter clipping to satisfy the Lipschitz condition in the algorithm induces a discontinuous vector field in the mean field dynamics, which gives rise to blow-up in finite time of the mean field dynamics. We look into a specific toy example that shows that all solutions to the mean field equations converge in the long time limit to time periodic solutions, this helps explain the failure to converge of the algorithm.
</description>
</item>

<item>
<title>
&#34;What is Different Between These Datasets?&#34; A Framework for Explaining Data Distribution Shifts
</title>
<link>
http://jmlr.org/papers/v26/24-0352.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0352/24-0352.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Varun Babbar*, Zhicheng Guo*, Cynthia Rudin</author>
<description>
The performance of machine learning models relies heavily on the quality of input data, yet real-world applications often face significant data-related challenges. A common issue arises when curating training data or deploying models: two datasets from the same domain may exhibit differing distributions. While many techniques exist for detecting such distribution shifts, there is a lack of comprehensive methods to explain these differences in a human-understandable way beyond opaque quantitative metrics. To bridge this gap, we propose a versatile framework of interpretable methods for comparing datasets. Using a variety of case studies, we demonstrate the effectiveness of our approach across diverse data modalities—including tabular data, text data, images, time-series signals – in both low and high-dimensional settings. These methods complement existing techniques by providing actionable and interpretable insights to better understand and address distribution shifts.
</description>
</item>

<item>
<title>
Assumption-lean and data-adaptive post-prediction inference
</title>
<link>
http://jmlr.org/papers/v26/24-0056.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0056/24-0056.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jiacheng Miao, Xinran Miao, Yixuan Wu, Jiwei Zhao, Qiongshi Lu</author>
<description>
A primary challenge facing modern scientific research is the limited availability of gold-standard data, which can be costly, labor-intensive, or invasive to obtain. With the rapid development of machine learning (ML), scientists can now employ ML algorithms to predict gold-standard outcomes using variables that are easier to obtain. However, these predicted outcomes are often used directly in subsequent statistical analyses, ignoring imprecision and heterogeneity introduced by the prediction procedure. This will likely result in false positive findings and invalid scientific conclusions. In this work, we introduce PoSt-Prediction Adaptive inference (PSPA) that allows valid and powerful inference based on ML-predicted data. Its “assumption-lean” property guarantees reliable statistical inference without assumptions on the ML prediction. Its “data-adaptive” feature guarantees an efficiency gain over existing methods, regardless of the accuracy of ML prediction. We demonstrate the statistical superiority and broad applicability of our method through simulations and real-data applications.
</description>
</item>

<item>
<title>
Bagged Regularized k-Distances for Anomaly Detection
</title>
<link>
http://jmlr.org/papers/v26/23-1519.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1519/23-1519.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yuchao Cai, Hanfang Yang, Yuheng Ma, Hanyuan Hang</author>
<description>
We consider the paradigm of unsupervised anomaly detection, which involves the identification of anomalies within a dataset in the absence of labeled examples. Though distance-based methods are top-performing for unsupervised anomaly detection, they suffer heavily from the sensitivity to the choice of the number of the nearest neighbors. In this paper, we propose a new distance-based algorithm called bagged regularized $k$-distances for anomaly detection (BRDAD), converting the unsupervised anomaly detection problem into a convex optimization problem. Our BRDAD algorithm selects the weights by minimizing the surrogate risk, i.e., the finite sample bound of the empirical risk of the bagged weighted $k$-distances for density estimation (BWDDE). This approach enables us to successfully address the sensitivity challenge of the hyperparameter choice in distance-based algorithms. Moreover, when dealing with large-scale datasets, the efficiency issues can be addressed by the incorporated bagging technique in our BRDAD algorithm. On the theoretical side, we establish fast convergence rates of the AUC regret of our algorithm and demonstrate that the bagging technique significantly reduces the computational complexity. On the practical side, we conduct numerical experiments to illustrate the insensitivity of the parameter selection of our algorithm compared with other state-of-the-art distance-based methods. Furthermore, our method achieves superior performance on real-world datasets with the introduced bagging technique compared to other approaches.
</description>
</item>

<item>
<title>
Four Axiomatic Characterizations of the Integrated Gradients Attribution Method
</title>
<link>
http://jmlr.org/papers/v26/23-0671.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0671/23-0671.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Daniel Lundstrom, Meisam Razaviyayn</author>
<description>
Deep neural networks have produced significant progress among machine learning models in terms of accuracy and functionality, but their inner workings are still largely unknown. Attribution  methods seek to shine  a light on these &#34;black box&#34; models by indicating how much each input contributed to a model&#39;s outputs. The Integrated Gradients (IG) method is a state of the art baseline attribution method in the axiomatic vein, meaning it is designed to conform to particular principles of attributions. We present four axiomatic characterizations of IG, establishing IG as the unique method satisfying four different sets of axioms.
</description>
</item>

<item>
<title>
Fast Algorithm for Constrained Linear Inverse Problems
</title>
<link>
http://jmlr.org/papers/v26/22-1380.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-1380/22-1380.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Mohammed Rayyan Sheriff, Floor Fenne Redel, Peyman Mohajerin Esfahani</author>
<description>
We consider the constrained Linear Inverse Problem (LIP), where a certain atomic norm (like the $\ell_1 $ norm) is minimized subject to a quadratic constraint. Typically, such cost functions are non-differentiable, which makes them not amenable to the fast optimization methods existing in practice. We propose two equivalent reformulations of the constrained LIP with improved convex regularity: (i) a smooth convex minimization problem, and (ii) a strongly convex min-max problem. These problems could be solved by applying existing acceleration-based convex optimization methods which provide better $ O \left( \frac{1}{k^2} \right)$ theoretical convergence guarantee, improving upon the current best rate of $O \left( \frac{1}{k} \right)$. We also provide a novel algorithm named the Fast Linear Inverse Problem Solver (FLIPS), which is tailored to maximally exploit the structure of the reformulations. We demonstrate the performance of FLIPS on the classical problems of Binary Selection, Compressed Sensing, and Image Denoising. We also provide open source \texttt{MATLAB} and \texttt{PYTHON} packages for these three examples, which can be easily adapted to other LIPs.
</description>
</item>

<item>
<title>
High-Rank Irreducible Cartesian Tensor Decomposition and Bases of Equivariant Spaces
</title>
<link>
http://jmlr.org/papers/v26/25-0134.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-0134/25-0134.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Shihao Shao, Yikang Li, Zhouchen Lin, Qinghua Cui</author>
<description>
Irreducible Cartesian tensors (ICTs) play a crucial role in the design of equivariant graph neural networks, as well as in theoretical chemistry and chemical physics. Meanwhile, the design space of available linear operations on tensors that preserve symmetry presents a significant challenge. The ICT decomposition and a basis of this equivariant space are difficult to obtain for high-rank tensors. After decades of research, Bonvicini (2024) has recently achieved an explicit ICT decomposition for $n=5$ with factorial time/space complexity. In this work we, for the first time, obtain decomposition matrices for ICTs up to rank $n=9$ with reduced and affordable complexity, by constructing what we call path matrices. The path matrices are obtained via performing chain-like contractions with Clebsch-Gordan matrices following the parentage scheme. We prove and leverage that the concatenation of path matrices is an orthonormal change-of-basis matrix between the Cartesian tensor product space and the spherical direct sum spaces. Furthermore, we identify a complete orthogonal basis for the equivariant space, rather than a spanning set (Pearce-Crump, 2023b), through this path matrices technique. Our method avoids the RREF algorithm and maintains a fully analytical derivation of each ICT decomposition matrix, thereby significantly improving the algorithm’s speed to obtain arbitrary rank orthogonal ICT decomposition matrices and orthogonal equivariant bases. We further extend our result to the arbitrary tensor product and direct sum spaces, enabling free design between different spaces while keeping symmetry. The Python code is available at https://github.com/ShihaoShao-GH/ICT-decomposition-and-equivariant-bases, where the $n=6,\dots,9$ ICT decomposition matrices are obtained in 1s, 3s, 11s, and 4m32s on 28-core Intel Xeon Gold 6330 CPU @ 2.00GHz, respectively.
</description>
</item>

<item>
<title>
Best Linear Unbiased Estimate from Privatized Contingency Tables
</title>
<link>
http://jmlr.org/papers/v26/24-1962.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1962/24-1962.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jordan Awan, Adam Edwards, Paul Bartholomew, Andrew Sillers</author>
<description>
In differential privacy (DP) mechanisms, it can be beneficial to release &#34;redundant&#34; outputs,  where some quantities can be estimated in multiple ways by  combining different privatized values. Indeed, the DP 2020 Decennial Census products published by the U.S. Census Bureau  consist of such redundant noisy counts. When redundancy is present, the DP output can be improved by enforcing self-consistency (i.e.,  estimators obtained using different noisy counts result in the same value), and we show that the minimum variance processing is a linear projection. However, standard projection algorithms require excessive computation and memory, making them impractical for large-scale applications such as the Decennial Census. We propose the Scalable Efficient Algorithm for Best Linear Unbiased Estimate (SEA BLUE), based on a two-step process of aggregation and differencing that 1) enforces self-consistency through a linear and unbiased procedure, 2) is computationally and memory efficient, 3) achieves the minimum variance solution under certain structural assumptions, and 4) is empirically shown to be robust to violations of these structural assumptions. We propose three methods of calculating confidence intervals from our estimates, under various assumptions. Finally, we apply SEA BLUE to two 2010 Census demonstration products, illustrating its scalability and validity.
</description>
</item>

<item>
<title>
Interpretable Global Minima of Deep ReLU Neural Networks on Sequentially Separable Data
</title>
<link>
http://jmlr.org/papers/v26/24-1516.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1516/24-1516.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Thomas Chen, Patrícia Muñoz Ewald</author>
<description>
We explicitly construct zero loss neural network classifiers.  We write the weight matrices and bias vectors in terms of cumulative parameters, which determine truncation maps acting recursively on input space. The configurations for the training data considered are $(i)$ sufficiently small, well separated clusters corresponding to each class, and $(ii)$ equivalence classes which are sequentially linearly separable. In the best case, for $Q$ classes of data in $\mathbb{R}^{M}$, global minimizers can be described with $Q(M+2)$ parameters.
</description>
</item>

<item>
<title>
Enhanced Feature Learning via Regularisation: Integrating Neural Networks and Kernel Methods
</title>
<link>
http://jmlr.org/papers/v26/24-1178.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1178/24-1178.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Bertille FOLLAIN, Francis BACH</author>
<description>
We propose a new method for feature learning and function estimation in supervised learning via regularised empirical risk minimisation. Our approach considers functions as expectations of Sobolev functions over all possible one-dimensional projections of the data. This framework is similar to kernel ridge regression, where the kernel is E_w(k(B)(wx, wx&#39;)), with k(B)(a, b) := min(|a|, |b|)1_{ab&gt;0} the Brownian kernel, and the distribution of the projections w is learnt. This can also be viewed as an infinite-width one-hidden layer neural network, optimising the first layer’s weights through gradient descent and explicitly adjusting the non-linearity and weights of the second layer. We introduce a gradient-based computational method for the estimator, called Brownian Kernel Neural Network (BKerNN), using particles to approximate the expectation, where the positive homogeneity of the Brownian kernel leads to improved robustness to local minima. Using Rademacher complexity, we show that BKerNN’s expected risk converges to the minimal risk with explicit high-probability rates of O(min((d/n)^1/2, n^−1/6)) (up to logarithmic factors).  Numerical experiments confirm our optimisation intuitions, and BKerNN outperforms kernel ridge regression, and favourably compares to a one-hidden layer neural network with ReLU activations in various settings and real datasets.
</description>
</item>

<item>
<title>
Data-Driven Performance Guarantees for Classical and Learned Optimizers
</title>
<link>
http://jmlr.org/papers/v26/24-0755.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0755/24-0755.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Rajiv Sambharya, Bartolomeo Stellato</author>
<description>
We introduce a data-driven approach to analyze the performance of continuous optimization algorithms using generalization guarantees from statistical learning theory. We study classical and learned optimizers to solve families of parametric optimization problems. We build generalization guarantees for classical optimizers, using a sample convergence bound, and for learned optimizers, using the Probably Approximately Correct (PAC)-Bayes framework. To train learned optimizers, we use a gradient-based algorithm to directly minimize the PAC-Bayes upper bound. Numerical experiments in signal processing, control, and meta-learning showcase the ability of our framework to provide strong generalization guarantees for both classical and learned optimizers given a fixed budget of iterations. For classical optimizers, our bounds which hold with high probability are much tighter than those that worst-case guarantees provide. For learned optimizers, our bounds outperform the empirical outcomes observed in their non-learned counterparts.
</description>
</item>

<item>
<title>
Contextual Bandits with Stage-wise Constraints
</title>
<link>
http://jmlr.org/papers/v26/24-0267.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0267/24-0267.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Aldo Pacchiano, Mohammad Ghavamzadeh, Peter Bartlett</author>
<description>
We study contextual bandits in the presence of a stage-wise constraint when the constraint must be satisfied both with high probability and in expectation. We start with the linear case where both the reward function and the stage-wise constraint (cost function) are linear. In each of the high probability and in expectation settings, we propose an upper-confidence bound algorithm for the problem and prove a $T$-round regret bound for it. We also prove a lower-bound for this constrained problem, show how our algorithms and analyses can be extended to multiple constraints, and provide simulations to validate our theoretical results. In the high probability setting, we describe the minimum requirements for the action set for our algorithm to be tractable. In the setting that the constraint is in expectation, we specialize our results to multi-armed bandits and propose a computationally efficient algorithm for this setting with regret analysis. Finally, we extend our results to the case where the reward and cost functions are both non-linear. We propose an algorithm for this case and prove a regret bound for it that characterize the function class complexity by the eluder dimension.
</description>
</item>

<item>
<title>
Boosting Causal Additive Models
</title>
<link>
http://jmlr.org/papers/v26/24-0052.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0052/24-0052.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Maximilian Kertel, Nadja Klein</author>
<description>
We present a boosting-based method to learn additive Structural Equation Models (SEMs) from observational data, with a focus on the theoretical aspects of determining the causal order among variables. We introduce a family of score functions based on arbitrary regression techniques, for which we establish sufficient conditions that guarantee consistent identification of the true causal ordering. Our analysis reveals that boosting with early stopping meets these criteria and thus offers a consistent score function for causal orderings.  To address the challenges posed by high-dimensional data sets, we adapt our approach through a component-wise gradient descent in the space of additive SEMs. Our simulation study supports the theoretical findings in low-dimensional settings and demonstrates that our high-dimensional adaptation is competitive with state-of-the-art methods. In addition, it exhibits robustness with respect to the choice of hyperparameters, thereby simplifying the tuning process.
</description>
</item>

<item>
<title>
Frequentist Guarantees of Distributed (Non)-Bayesian Inference
</title>
<link>
http://jmlr.org/papers/v26/23-1504.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1504/23-1504.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Bohan Wu, César A. Uribe</author>
<description>
We establish frequentist properties, i.e., posterior consistency, asymptotic normality, and posterior contraction rates, for the distributed (non-)Bayesian inference problem for a set of agents connected over a network. These results are motivated by the need to analyze large, decentralized datasets, where distributed (non)-Bayesian inference has become a critical research area across multiple fields, including statistics, machine learning, and economics. Our results show that, under appropriate assumptions on the communication graph, distributed (non)-Bayesian inference retains parametric efficiency while enhancing robustness in uncertainty quantification.  We also explore the trade-off between statistical efficiency and communication efficiency by examining how the design and size of the communication graph impact the posterior contraction rate. Furthermore, we extend our analysis to time-varying graphs and apply our results to exponential family models, distributed logistic regression, and decentralized detection models.
</description>
</item>

<item>
<title>
Asymptotic Inference for Multi-Stage Stationary Treatment Policy with Variable Selection
</title>
<link>
http://jmlr.org/papers/v26/23-0660.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0660/23-0660.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Daiqi Gao, Yufeng Liu, Donglin Zeng</author>
<description>
Dynamic treatment regimes or policies are a sequence of decision functions over multiple stages that are tailored to individual features. One important class of treatment policies in practice, namely multi-stage stationary treatment policies, prescribes treatment assignment probabilities using the same decision function across stages, where the decision is based on the same set of features consisting of time-evolving variables (e.g., routinely collected disease biomarkers). Although there has been extensive literature on constructing valid inference for the value function associated with dynamic treatment policies, little work has focused on the policies themselves, especially in the presence of high-dimensional features. We aim to fill the gap in this work. Specifically, we first obtain the multi-stage stationary treatment policy by minimizing the negative augmented inverse probability weighted estimator of the value function to increase asymptotic efficiency. An $L_1$ penalty is applied on the policy parameters to select important features. We then construct one-step improvements of the policy parameter estimators for valid inference. Theoretically, we show that the improved estimators are asymptotically normal, even if nuisance parameters are estimated at a slow convergence rate and the dimension of the features increases with the sample size. Our numerical studies demonstrate that the proposed method estimates a sparse policy with a near-optimal value function and conducts valid inference for the policy parameters.
</description>
</item>

<item>
<title>
EMaP: Explainable AI with Manifold-based Perturbations
</title>
<link>
http://jmlr.org/papers/v26/22-1157.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-1157/22-1157.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Minh Nhat Vu, Huy Quang Mai, My T. Thai</author>
<description>
In the last few years, many explanation methods based on the perturbations of input data have been introduced to shed light on the predictions generated by black-box models. The goal of this work is to introduce a novel perturbation scheme so that more faithful and robust explanations can be obtained. Our study focuses on the impact of perturbing directions on the data topology. We show that perturbing along the orthogonal directions of the input manifold better preserves the data topology, both in the worst-case analysis of the discrete Gromov-Hausdorff distance and in the average-case analysis via persistent homology. From those results, we introduce EMaP algorithm, realizing the orthogonal perturbation scheme. Our experiments show that EMaP not only improves the explainers&#39; performance but also helps them overcome a recently developed attack against perturbation-based explanation methods.
</description>
</item>

<item>
<title>
Autoencoders in Function Space
</title>
<link>
http://jmlr.org/papers/v26/25-0035.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-0035/25-0035.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Justin Bunker, Mark Girolami, Hefin Lambley, Andrew M. Stuart, T. J. Sullivan</author>
<description>
Autoencoders have found widespread application in both their original deterministic form and in their variational formulation (VAEs). In scientific applications and in image processing it is often of interest to consider data that are viewed as functions; while discretisation (of differential equations arising in the sciences) or pixellation (of images) renders problems finite dimensional in practice, conceiving first of algorithms that operate on functions, and only then discretising or pixellating, leads to better algorithms that smoothly operate between resolutions. In this paper function-space versions of the autoencoder (FAE) and variational autoencoder (FVAE) are introduced, analysed, and deployed. Well-definedness of the objective governing VAEs is a subtle issue, particularly in function space, limiting applicability. For the FVAE objective to be well defined requires compatibility of the data distribution with the chosen generative model; this can be achieved, for example, when the data arise from a stochastic differential equation, but is generally restrictive. The FAE objective, on the other hand, is well defined in many situations where FVAE fails to be. Pairing the FVAE and FAE objectives with neural operator architectures that can be evaluated on any mesh enables new applications of autoencoders to inpainting, superresolution, and generative modelling of scientific data.
</description>
</item>

<item>
<title>
Nonparametric Regression on Random Geometric Graphs Sampled from Submanifolds
</title>
<link>
http://jmlr.org/papers/v26/24-1960.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1960/24-1960.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Paul Rosa, Judith Rousseau</author>
<description>
We consider the nonparametric regression problem when the covariates are located on an unknown compact submanifold of a Euclidean space. Under defining a random geometric graph structure over the covariates we analyse the asymptotic frequentist behaviour of the posterior distribution arising from Bayesian priors designed through random basis expansion in the graph Laplacian eigenbasis. Under Hölder smoothness assumption on the regression function and the density of the covariates over the submanifold, we prove that the posterior contraction rates of such methods are minimax optimal (up to logarithmic factors) for any positive smoothness index.
</description>
</item>

<item>
<title>
System Neural Diversity: Measuring Behavioral Heterogeneity in Multi-Agent Learning
</title>
<link>
http://jmlr.org/papers/v26/24-1477.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1477/24-1477.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Matteo Bettini, Ajay Shankar, Amanda Prorok</author>
<description>
Evolutionary science provides evidence that diversity confers resilience in natural systems. Yet, traditional multi-agent reinforcement learning techniques commonly enforce homogeneity to increase training sample efficiency. When a system of learning agents is not constrained to homogeneous policies, individuals may develop diverse behaviors, resulting in emergent complementarity that benefits the system. Despite this, there is a surprising lack of tools that quantify behavioral diversity. Such techniques would pave the way towards understanding the impact of diversity in collective artificial intelligence and enabling its control. In this paper, we introduce System Neural Diversity (SND): a measure of behavioral heterogeneity in multi-agent systems. We discuss and prove its theoretical properties, and compare it with alternate, state-of-the-art behavioral diversity metrics used in the robotics domain. Through simulations of a variety of cooperative multi-robot tasks, we show how our metric constitutes an important tool that enables measurement and control of behavioral heterogeneity. In dynamic tasks, where the problem is affected by repeated disturbances during training, we show that SND allows us to measure latent resilience skills acquired by the agents, while other proxies, such as task performance (reward), fail to. Finally, we show how the metric can be employed to control diversity, allowing us to enforce a desired heterogeneity set-point or range. We demonstrate how this paradigm can be used to bootstrap the exploration phase, finding optimal policies faster, thus enabling novel and more efficient MARL paradigms.
</description>
</item>

<item>
<title>
Distribution Estimation under the Infinity Norm
</title>
<link>
http://jmlr.org/papers/v26/24-1166.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1166/24-1166.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Aryeh Kontorovich, Amichai Painsky</author>
<description>
We present novel bounds for estimating discrete probability distributions under the $\ell_\infty$ norm. These are nearly optimal in various precise senses, including a kind of instance-optimality. Our data-dependent convergence guarantees for the maximum likelihood estimator significantly improve upon the currently known results. A variety of techniques are utilized and innovated upon, including Chernoff-type inequalities and empirical Bernstein bounds. We illustrate our results in synthetic and real-world experiments. Finally, we apply our proposed framework to a basic selective inference problem, where we estimate the most frequent probabilities in a sample.
</description>
</item>

<item>
<title>
Extending Temperature Scaling with Homogenizing Maps
</title>
<link>
http://jmlr.org/papers/v26/24-0700.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0700/24-0700.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Christopher Qian, Feng Liang, Jason Adams</author>
<description>
As machine learning models continue to grow more complex, poor calibration significantly limits the reliability of their predictions. Temperature scaling learns a single temperature parameter to scale the output logits, and despite its simplicity, remains one of the most effective post-hoc recalibration methods. We identify one of temperature scaling&#39;s defining attributes, that it increases the uncertainty of the predictions in a manner that we term homogenization, and propose to learn the optimal recalibration mapping from a larger class of functions that satisfies this property. We demonstrate the advantage of our method over temperature scaling in both calibration and out-of-distribution detection. Additionally, we extend our methodology and experimental evaluation to recalibration in the Bayesian setting.
</description>
</item>

<item>
<title>
Density Estimation Using the Perceptron
</title>
<link>
http://jmlr.org/papers/v26/24-0261.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0261/24-0261.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Patrik Róbert Gerber, Tianze Jiang, Yury Polyanskiy, Rui Sun</author>
<description>
We propose a new density estimation algorithm. Given
$n$ i.i.d. observations from a distribution belonging to a class
of densities on $\mathbb{R}^d$, our estimator outputs any density in the class whose “perceptron
discrepancy” with the empirical distribution is at most $O(\sqrt{d/n})$.
The perceptron discrepancy is defined as the largest
difference in mass two distribution place on any halfspace. It is shown that
this estimator achieves the expected total variation distance to the truth that is almost
minimax optimal over the class of densities with bounded Sobolev norm and Gaussian
mixtures. This suggests that the regularity of the prior distribution could be an
explanation for the efficiency of the ubiquitous step in machine learning that replaces optimization over large function spaces with simpler parametric
classes (such as discriminators of GANs).
We also show that replacing the perceptron discrepancy with
the generalized energy distance of  Székely and Rizzo (2013) further improves
total variation loss. The generalized energy distance between empirical
distributions is easily computable and
differentiable, which makes it especially useful for fitting generative models.
To the best of our knowledge, it is the first “simple” distance with such
properties that yields minimax optimal statistical guarantees. 


In addition, we shed light on the ubiquitous method of representing discrete data in domain $[k]$ via embedding vectors on a unit ball in $\mathbb{R}^d$. We show that taking $d \asymp \log(k)$ allows one to use simple linear probing to evaluate and estimate total variation distance, as well as recovering minimax optimal sample complexity for the class of discrete distributions on $[k]$.
</description>
</item>

<item>
<title>
Simplex Constrained Sparse Optimization via Tail Screening
</title>
<link>
http://jmlr.org/papers/v26/24-0010.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0010/24-0010.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Peng Chen, Jin Zhu, Junxian Zhu, Xueqin Wang</author>
<description>
We consider the probabilistic simplex-constrained sparse recovery problem. The commonly used Lasso-type penalty for promoting sparsity is ineffective in this context since it is a constant within the simplex. Despite this challenge, fortunately, simplex constraint itself brings a self-regularization property, i.e., the empirical risk minimizer without any sparsity-promoting procedure obtains the usual Lasso-type estimation error. Moreover, we analyze the iterates of a projected gradient descent method and show its convergence to the ground truth sparse solution in the geometric rate until a satisfied statistical precision is attained. Although the estimation error is statistically optimal, the resulting solution is usually more dense than the sparse ground truth. To further sparsify the iterates, we propose a method called PERMITS via embedding a tail screening procedure, i.e., identifying negligible components and discarding them during iterations, into the projected gradient descent method. Furthermore, we combine tail screening and the special information criterion to balance the trade-off between fitness and complexity. Theoretically, the proposed PERMITS method can exactly recover the ground truth support set under mild conditions and thus obtain the oracle property. We demonstrate the statistical and computational efficiency of PERMITS with both synthetic and real data. The implementation of the proposed method can be found in https://github.com/abess-team/PERMITS.
</description>
</item>

<item>
<title>
Score-Based Diffusion Models in Function Space
</title>
<link>
http://jmlr.org/papers/v26/23-1472.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1472/23-1472.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jae Hyun Lim, Nikola B. Kovachki, Ricardo Baptista, Christopher Beckham, Kamyar Azizzadenesheli, Jean Kossaifi, Vikram Voleti, Jiaming Song, Karsten Kreis, Jan Kautz, Christopher Pal, Arash Vahdat, Anima Anandkumar</author>
<description>
Diffusion models have recently emerged as a powerful framework for generative modeling. They consist of a forward process that perturbs input data with Gaussian white noise and a reverse process that learns a score function to generate samples by denoising. Despite their tremendous success, they are mostly formulated on finite-dimensional spaces, e.g., Euclidean, limiting their applications to many domains where the data has a functional form, such as in scientific computing and 3D geometric data analysis. This work introduces a mathematically rigorous framework called Denoising Diffusion Operators (DDOs) for training diffusion models in function space. In DDOs, the forward process perturbs input functions gradually using a Gaussian process. The generative process is formulated by a function-valued annealed Langevin dynamic. Our approach requires an appropriate notion of the score for the perturbed data distribution, which we obtain by generalizing denoising score matching to function spaces that can be infinite-dimensional. We show that the corresponding discretized algorithm generates accurate samples at a fixed cost independent of the data resolution. We theoretically and numerically verify the applicability of our approach on a set of function-valued problems, including generating solutions to the Navier-Stokes equation viewed as the push-forward distribution of forcings from a Gaussian Random Field (GRF), as well as volcano InSAR and MNIST-SDF.
</description>
</item>

<item>
<title>
Regularized Rényi Divergence Minimization through Bregman Proximal Gradient Algorithms
</title>
<link>
http://jmlr.org/papers/v26/23-0573.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0573/23-0573.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Thomas Guilmeau, Emilie Chouzenoux, Víctor Elvira</author>
<description>
We study the variational inference problem of minimizing a regularized Rényi divergence over an exponential family. We propose to solve this problem with a Bregman proximal gradient algorithm. We propose a sampling-based algorithm to cover the black-box setting, corresponding to a stochastic Bregman proximal gradient algorithm with biased gradient estimator. We show that the resulting algorithms can be seen as relaxed moment-matching algorithms with an additional proximal step. Using Bregman updates instead of Euclidean ones allows us to exploit the geometry of our approximate model. We prove strong convergence guarantees for both our deterministic and stochastic algorithms using this viewpoint, including monotonic decrease of the objective, convergence to a stationary point or to the minimizer, and geometric convergence rates. These new theoretical insights lead to a versatile, robust, and competitive method, as illustrated by numerical experiments
</description>
</item>

<item>
<title>
WEFE: A Python Library for Measuring and Mitigating Bias in Word Embeddings
</title>
<link>
http://jmlr.org/papers/v26/22-1133.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-1133/22-1133.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Pablo Badilla, Felipe Bravo-Marquez, María José Zambrano, Jorge Pérez</author>
<description>
Word embeddings, which are a mapping of words into continuous vectors, are widely used in modern Natural Language Processing (NLP) systems. However, they are prone to inherit stereotypical social biases from the corpus on which they are built. 
The research community has focused on two main tasks to address this problem: 1) how to measure these biases, and 2) how to mitigate them.
Word Embedding Fairness Evaluation (WEFE) is an open source library that implements many fairness metrics and mitigation methods in a unified framework. It also provides a standard interface for designing new ones. 
The software follows the object-oriented paradigm with a strong focus on extensibility. Each of its methods is appropriately documented, verified and tested.
WEFE is not limited to just a library: it also contains several replications of previous studies as well as tutorials that serve as educational material for newcomers to the field.
It is licensed under BSD-3 and can be easily installed through pip and conda package managers.
</description>
</item>

<item>
<title>
Frontiers to the learning of nonparametric hidden Markov models
</title>
<link>
http://jmlr.org/papers/v26/24-2230.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-2230/24-2230.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Kweku Abraham, Elisabeth Gassiat, Zacharie Naulet</author>
<description>
Hidden Markov models (HMMs) are flexible tools for clustering dependent data coming from unknown populations, allowing nonparametric modelling of the population densities. Identifiability fails when the data is in fact independent and identically distributed (i.i.d.), and we study the frontier between learnable and unlearnable two-state nonparametric HMMs. Learning the parameters of the HMM requires solving a nonlinear inverse problem whose difficulty depends not only on the smoothnesses of the populations but also on the distance to the i.i.d. boundary of the parameter set. The latter difficulty is mostly ignored in the literature in favour of assumptions precluding nearly independent data. This is the first work conducting a precise nonasymptotic, nonparametric analysis of the minimax risk taking into account all aspects of the hardness of the problem, in the case of two populations. Our analysis reveals an unexpected interplay between the distance to the i.i.d. boundary and the relative smoothnesses of the two populations: a surprising and intriguing transition occurs in the rate when the two densities have differing smoothnesses. We obtain upper and lower bounds revealing that, close to the i.i.d. boundary, it is possible to &#34;borrow strength&#34; from the estimator of the smoother density to improve the risk of the other.
</description>
</item>

<item>
<title>
On Non-asymptotic Theory of Recurrent Neural Networks in Temporal Point Processes
</title>
<link>
http://jmlr.org/papers/v26/24-1953.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1953/24-1953.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Zhiheng Chen, Guanhua Fang, Wen Yu</author>
<description>
Temporal point process (TPP) is an important tool for modeling and predicting irregularly timed events across various domains. Recently, the recurrent neural network (RNN)-based TPPs have shown practical advantages over traditional parametric TPP models. However, in the current literature, it remains nascent in understanding neural TPPs from theoretical viewpoints. In this paper, we establish the excess risk bounds of RNN-TPPs under many well-known TPP settings. We especially show that an RNN-TPP with no more than four layers can achieve vanishing generalization errors. Our technical contributions include the characterization of the complexity of the multi-layer RNN class, the construction of $\tanh$ neural networks for approximating dynamic event intensity functions, and the truncation technique for alleviating the issue of unbounded event sequences. Our results bridge the gap between TPP&#39;s application and neural network theory.
</description>
</item>

<item>
<title>
Classification in the high dimensional Anisotropic mixture framework: A new take on Robust Interpolation
</title>
<link>
http://jmlr.org/papers/v26/24-1366.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1366/24-1366.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Stanislav Minsker, Mohamed Ndaoud, Yiqiu Shen</author>
<description>
We study the classification problem under the two-component anisotropic sub-Gaussian mixture model in high dimensions and in the non-asymptotic setting. First, we derive lower bounds and  matching upper bounds for the minimax risk of classification in this framework. We also show that in the high-dimensional regime, the linear discriminant analysis classifier turns out to be sub-optimal in the minimax sense. Next, we give precise characterization of the risk of classifiers based on solutions of $\ell_2$-regularized least squares problem. We deduce that the interpolating solutions may outperform the regularized classifiers under mild assumptions on the covariance structure of the noise, and present concrete examples of this phenomenon. Our analysis also demonstrates robustness of interpolation to certain models of corruption. To the best of our knowledge, this peculiar fact has not yet been investigated in the rapidly growing literature related to interpolation. We conclude that interpolation is not only benign but can also be optimal, and in some cases robust.
</description>
</item>

<item>
<title>
Universal Online Convex Optimization Meets Second-order Bounds
</title>
<link>
http://jmlr.org/papers/v26/24-1131.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1131/24-1131.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Lijun Zhang, Yibo Wang, Guanghui Wang, Jinfeng Yi, Tianbao Yang</author>
<description>
Recently, several universal methods have been proposed for online convex optimization, and attain minimax rates for multiple types of convex  functions simultaneously. However, they need to design and optimize one surrogate loss for each type of functions, making it difficult to exploit the structure of the problem and utilize existing algorithms. In this paper, we propose a simple strategy for universal online convex optimization, which avoids these limitations. The key idea is to construct a set of experts to process the original online functions, and deploy a meta-algorithm over the linearized losses to aggregate predictions from experts. Specifically, the meta-algorithm is required to yield a second-order bound with excess losses, so that it can leverage strong convexity and exponential concavity to control the meta-regret. In this way, our strategy inherits the theoretical guarantee of any expert designed for strongly convex functions and exponentially concave functions, up to a double logarithmic factor. As a result, we can plug in off-the-shelf online solvers as black-box experts to deliver problem-dependent regret bounds. For general convex functions, it maintains the minimax optimality and also achieves a small-loss bound. Furthermore, we extend our universal strategy to online composite optimization, where the loss function comprises a time-varying function and a fixed regularizer. To deal with the composite loss functions, we employ a meta-algorithm based on the optimistic online learning framework, which not only enjoys a second-order bound, but also can utilize estimations for upcoming loss functions.  With suitable configurations, we show that the additional regularizer does not contribute to the meta-regret, thus ensuring the universality in the composite setting.
</description>
</item>

<item>
<title>
Sample Complexity of the Linear Quadratic Regulator: A Reinforcement Learning Lens
</title>
<link>
http://jmlr.org/papers/v26/24-0636.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0636/24-0636.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Amirreza Neshaei Moghaddam, Alex Olshevsky, Bahman Gharesifard</author>
<description>
We provide the first known algorithm that provably achieves $\varepsilon$-optimality within $\widetilde{O}(1/\varepsilon)$ function evaluations for the discounted discrete-time linear quadratic regulator problem with unknown parameters, without relying on two-point gradient estimates. These estimates are known to be unrealistic in many settings, as they depend on using the exact same initialization, which is to be selected randomly, for two different policies. Our results substantially improve upon the existing literature outside the realm of two-point gradient estimates, which either leads to $\widetilde{O}(1/\varepsilon^2)$ rates or heavily relies on stability assumptions.
</description>
</item>

<item>
<title>
Randomization Can Reduce Both Bias and Variance: A Case Study in Random Forests
</title>
<link>
http://jmlr.org/papers/v26/24-0255.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0255/24-0255.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Brian Liu, Rahul Mazumder</author>
<description>
We study the often overlooked phenomenon, first noted in Breiman (2001), that random forests appear to reduce bias compared to bagging. Motivated by an interesting paper by Mentch and Zhou (2020), where the authors explain the success of random forests in low signal-to-noise ratio (SNR) settings through regularization, we explore how random forests can capture patterns in the data that bagging ensembles fail to capture. We empirically demonstrate that in the presence of such patterns, random forests reduce bias along with variance and can increasingly outperform bagging ensembles when SNR is high. Our observations offer insights into the real-world success of random forests across a range of SNRs and enhance our understanding of the difference between random forests and bagging ensembles. Our investigations also yield practical insights into the importance of tuning $mtry$ in random forests.
</description>
</item>

<item>
<title>
skglm: Improving scikit-learn for Regularized Generalized Linear Models
</title>
<link>
http://jmlr.org/papers/v26/24-0008.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0008/24-0008.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Badr Moufad, Pierre-Antoine Bannier, Quentin Bertrand, Quentin Klopfenstein, Mathurin Massias</author>
<description>
We introduce skglm, an open-source Python package for regularized Generalized Linear Models. Thanks to its composable nature, it supports combining datafits, penalties, and solvers to fit a wide range of models, many of them not included in scikit-learn (e.g. Group Lasso and variants). It uses state-of-the-art algorithms to solve problems involving high-dimensional datasets, providing large speed-ups compared to existing implementations. It is fully compliant with the scikit-learn API and acts as a drop-in replacement for its estimators. Finally, it abides by the standards of open source development and is integrated in the scikit-learn-contrib GitHub organization.
</description>
</item>

<item>
<title>
Losing Momentum in Continuous-time Stochastic Optimisation
</title>
<link>
http://jmlr.org/papers/v26/23-1396.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1396/23-1396.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Kexin Jin, Jonas Latz, Chenguang Liu, Alessandro Scagliotti</author>
<description>
The training of modern machine learning models often consists in solving high-dimensional non-convex optimisation problems that are subject to large-scale data. In this context, momentum-based stochastic optimisation algorithms have become particularly widespread. The stochasticity arises from data subsampling which reduces computational cost. Both, momentum and stochasticity help the algorithm to converge globally. In this work, we propose and analyse a continuous-time model for stochastic gradient descent with momentum. This model is a piecewise-deterministic Markov process that represents the optimiser by an underdamped dynamical system and the data subsampling through a stochastic switching. We investigate longtime limits, the subsampling-to-no-subsampling limit, and the momentum-to-no-momentum limit. We are particularly interested in the case of reducing the momentum over time. Under convexity assumptions, we show convergence of our dynamical system to the global minimiser when reducing momentum over time and letting the subsampling rate go to infinity. We then propose a stable, symplectic discretisation scheme to construct an algorithm from our continuous-time dynamical system. In experiments, we study our scheme in convex and non-convex test problems. Additionally, we train a convolutional neural network in an image classification problem. Our algorithm attains competitive results compared to stochastic gradient descent with momentum.
</description>
</item>

<item>
<title>
Latent Process Models for Functional Network Data
</title>
<link>
http://jmlr.org/papers/v26/23-0444.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0444/23-0444.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Peter W. MacDonald, Elizaveta Levina, Ji Zhu</author>
<description>
Network data are often sampled with auxiliary information or collected through the observation of a complex system over time, leading to multiple network snapshots indexed by a continuous variable. Many methods in statistical network analysis are traditionally designed for a single network, and can be applied to an aggregated network in this setting, but that approach can miss important functional structure. Here we develop an approach to estimating the expected network explicitly as a function of a continuous index, be it time or another indexing variable. We parameterize the network expectation through low dimensional latent processes, whose components we represent with a fixed, finite-dimensional functional basis. We derive a gradient descent estimation algorithm, establish theoretical guarantees for recovery of the low dimensional structure, compare our method to competitors, and apply it to a data set of international political interactions over time, showing our proposed method to adapt well to data, outperform competitors, and provide interpretable and meaningful results.
</description>
</item>

<item>
<title>
Dynamic Bayesian Learning for Spatiotemporal Mechanistic Models
</title>
<link>
http://jmlr.org/papers/v26/22-0896.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0896/22-0896.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Sudipto Banerjee, Xiang Chen, Ian Frankenburg, Daniel Zhou</author>
<description>
We develop an approach for Bayesian learning of spatiotemporal dynamical mechanistic models. Such learning consists of statistical emulation of the mechanistic system that can efficiently interpolate the output of the system from arbitrary inputs. The emulated learner can then be used to train the system from noisy data achieved by melding information from observed data with the emulated mechanistic system. This joint melding of mechanistic systems employ hierarchical state-space models with Gaussian process regression. Assuming the dynamical system is controlled by a finite collection of inputs, Gaussian process regression learns the effect of these parameters through a number of training runs, driving the stochastic innovations of the spatiotemporal state-space component. This enables efficient modeling of the dynamics over space and time. This article details exact inference with analytically accessible posterior distributions in hierarchical matrix-variate Normal and Wishart models in designing the emulator. This step obviates expensive iterative algorithms such as Markov chain Monte Carlo or variational approximations. We also show how emulation is applicable to large-scale emulation by designing a dynamic Bayesian transfer learning framework. Inference on mechanistic model parameters proceeds using Markov chain Monte Carlo as a post-emulation step using the emulator as a regression component. We demonstrate this framework through solving inverse problems arising in the analysis of ordinary and partial nonlinear differential equations and, in addition, to a black-box computer model generating spatiotemporal dynamics across a graphical model.
</description>
</item>

<item>
<title>
On the Ability of Deep Networks to Learn Symmetries from Data: A Neural Kernel Theory
</title>
<link>
http://jmlr.org/papers/v26/24-2175.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-2175/24-2175.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Andrea Perin, Stephane Deny</author>
<description>
Symmetries (transformations by group actions) are present in many datasets, and leveraging them holds considerable promise for improving predictions in machine learning. In this work, we aim to understand when and how deep networks---with standard architectures trained in a standard, supervised way---learn symmetries from data. Inspired by real-world scenarios, we study a classification paradigm where data symmetries are only partially observed during training: some classes include all transformations of a cyclic group, while others---only a subset. We ask: under which conditions will deep networks correctly classify the partially sampled classes?
In the infinite-width limit, where neural networks behave like kernel machines, we derive a neural kernel theory of symmetry learning. The group-cyclic nature of the dataset allows us to analyze the Gram matrix of neural kernels in the Fourier domain; here we find a simple characterization of the generalization error as a function of class separation (signal) and class-orbit density (noise). This characterization reveals that generalization can only be successful when the local structure of the data prevails over its non-local, symmetry-induced structure, in the kernel space defined by the architecture. This occurs when (1) classes are sufficiently distinct and (2) class orbits are sufficiently dense.
We extend our theoretical treatment to any finite group, including non-abelian groups. Our framework also applies to equivariant architectures (e.g., CNNs), and recovers their success in the special case where the architecture matches the inherent symmetry of the data. Empirically, our theory reproduces the generalization failure of finite-width networks (MLP, CNN, ViT) trained on partially observed versions of rotated-MNIST. We conclude that conventional deep networks lack a mechanism to learn symmetries that have not been explicitly embedded in their architecture a priori. In the future, our framework could be extended to guide the design of architectures and training procedures able to learn symmetries from data.
</description>
</item>

<item>
<title>
Fine-grained Analysis and Faster Algorithms for Iteratively Solving Linear Systems
</title>
<link>
http://jmlr.org/papers/v26/24-1906.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1906/24-1906.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Michal Dereziński, Daniel LeJeune, Deanna Needell, Elizaveta Rebrova</author>
<description>
Despite being a key bottleneck in many machine learning tasks, the cost of solving large linear systems has proven challenging to quantify due to problem-dependent quantities such as condition numbers.
To tackle this, we
    consider a fine-grained notion of complexity for solving linear systems, which is motivated by applications where the data exhibits low-dimensional structure, including spiked covariance models and kernel machines, and when the linear system is explicitly regularized, such as ridge regression.
    
    Concretely, let $\kappa_\ell$ be the ratio between the $\ell$th largest and the smallest singular value of $n\times n$ matrix $A$.
    We give a stochastic algorithm based on the Sketch-and-Project paradigm, that solves the linear system $Ax=b$ in time
    $\tilde O(\kappa_\ell\cdot n^2\log1/\epsilon)$ for any $\ell = O(n^{0.729})$.
This is a direct improvement over 
preconditioned conjugate gradient, and it provides a stronger separation between stochastic linear solvers and algorithms accessing $A$ only through matrix-vector products.

Our main technical contribution is the new analysis of the first and second moments of the random projection matrix that arises in Sketch-and-Project.
</description>
</item>

<item>
<title>
Deep Generative Models: Complexity, Dimensionality, and Approximation
</title>
<link>
http://jmlr.org/papers/v26/24-1335.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1335/24-1335.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Kevin Wang, Hongqian Niu, Yixin Wang, Didong Li</author>
<description>
Generative networks have shown remarkable success in learning complex data distributions, particularly in generating high-dimensional data from lower-dimensional inputs. While this capability is well-documented empirically, its theoretical underpinning remains unclear. One common theoretical explanation appeals to the widely accepted manifold hypothesis, which suggests that many real-world datasets, such as images and signals, often possess intrinsic low-dimensional geometric structures. Under this manifold hypothesis, it is widely believed that to approximate a distribution on a $d$-dimensional Riemannian manifold, the latent dimension needs to be at least $d$ or $d+1$. In this work, we show that this requirement on the latent dimension is not necessary by demonstrating that generative networks can approximate distributions on $d$-dimensional Riemannian manifolds from inputs of any arbitrary dimension, even lower than $d$, taking inspiration from the concept of space-filling curves. This approach, in turn, leads to a super-exponential complexity bound of the deep neural networks through expanded neurons. Our findings thus challenge the conventional belief on the relationship between input dimensionality and the ability of generative networks to model data distributions. This novel insight not only corroborates the practical effectiveness of generative networks in handling complex data structures, but also underscores a critical trade-off between approximation error, dimensionality, and model complexity.
</description>
</item>

<item>
<title>
ClimSim-Online: A Large Multi-Scale Dataset and Framework for Hybrid Physics-ML Climate Emulation
</title>
<link>
http://jmlr.org/papers/v26/24-1014.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1014/24-1014.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Sungduk Yu, Zeyuan Hu, Akshay Subramaniam, Walter Hannah, Liran Peng, Jerry Lin, Mohamed Aziz Bhouri, Ritwik Gupta, Björn Lütjens, Justus C. Will, Gunnar Behrens, Julius J. M. Busecke, Nora Loose, Charles I Stern, Tom Beucler, Bryce Harrop, Helge Heuer, Benjamin R Hillman, Andrea Jenney, Nana Liu, Alistair White, Tian Zheng, Zhiming Kuang, Fiaz Ahmed, Elizabeth Barnes, Noah D. Brenowitz, Christopher Bretherton, Veronika Eyring, Savannah Ferretti, Nicholas Lutsko, Pierre Gentine, Stephan Mandt, J. David Neelin, Rose Yu, Laure Zanna, Nathan M. Urban, Janni Yuval, Ryan Abernathey, Pierre Baldi, Wayne Chuang, Yu Huang, Fernando Iglesias-Suarez, Sanket Jantre, Po-Lun Ma, Sara Shamekh, Guang Zhang, Michael Pritchard</author>
<description>
Modern climate projections lack adequate spatial and temporal resolution due to computational constraints, leading to inaccuracies in representing critical processes like thunderstorms that occur on the sub-resolution scale. Hybrid methods combining physics with machine learning (ML) offer faster, higher fidelity climate simulations by outsourcing compute-hungry, high-resolution simulations to ML emulators. However, these hybrid physics-ML simulations require domain-specific data and workflows that have been inaccessible to many ML experts. This paper is an extended version of our NeurIPS award-winning ClimSim dataset paper. The ClimSim dataset includes 5.7 billion pairs of multivariate input/output vectors spanning ten years at high temporal resolution, capturing the influence of high-resolution, high-fidelity physics on a host climate simulator&#39;s macro-scale state. In this extended version, we introduce a significant new contribution in Section 5, which provides a cross-platform, containerized pipeline to integrate ML models into operational climate simulators for hybrid testing. We also implement various baselines of ML models and hybrid simulators to highlight the ML challenges of building stable, skillful emulators. The data (https://huggingface.co/datasets/LEAP/ClimSim_high-res, also in a low-resolution version at https://huggingface.co/datasets/LEAP/ClimSim_low-res and an aquaplanet version at https://huggingface.co/datasets/LEAP/ClimSim_low-res_aqua-planet) and code (https://leap-stc.github.io/ClimSim and https://github.com/leap-stc/climsim-online) are publicly released to support the development of hybrid physics-ML and high-fidelity climate simulations.
</description>
</item>

<item>
<title>
Conditional Wasserstein Distances with Applications in Bayesian OT Flow Matching
</title>
<link>
http://jmlr.org/papers/v26/24-0586.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0586/24-0586.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jannis Chemseddine, Paul Hagemann, Gabriele Steidl, Christian Wald</author>
<description>
In inverse problems, many conditional generative models approximate the posterior measure by minimizing a distance between the joint measure and its learned approximation. While this approach also controls the distance between the posterior measures in the case of the Kullback–Leibler divergence, the same in general does not hold true for the Wasserstein distance. In this paper, we introduce a conditional Wasserstein distance via a set of restricted couplings that equals the expected Wasserstein distance of the posteriors. Interestingly, the dual formulation of the conditional Wasserstein-1 distance resembles losses in the conditional Wasserstein GAN literature in a quite natural way. We derive theoretical properties of the conditional Wasserstein distance, characterize the corresponding geodesics and velocity fields as well as the flow ODEs. Subsequently, we propose to approximate the velocity fields by relaxing the conditional Wasserstein distance. Based on this, we propose an extension of OT Flow Matching for solving Bayesian inverse problems and demonstrate its numerical advantages on an inverse problem and class-conditional image generation.
</description>
</item>

<item>
<title>
Deep Variational Multivariate Information Bottleneck - A Framework for Variational Losses
</title>
<link>
http://jmlr.org/papers/v26/24-0204.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0204/24-0204.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Eslam Abdelaleem, Ilya Nemenman, K. Michael Martini</author>
<description>
Variational dimensionality reduction methods are widely used for their accuracy, generative capabilities, and robustness. We introduce a unifying framework that generalizes both such as traditional and state-of-the-art methods. The framework is based on an interpretation of the multivariate information bottleneck, trading off the information preserved in an encoder graph (defining what to compress) against that in a decoder graph (defining a generative model for data). Using this approach, we rederive existing methods, including the deep variational information bottleneck, variational autoencoders, and deep multiview information bottleneck. We naturally extend the deep variational CCA (DVCCA) family to beta-DVCCA and introduce a new method, the deep variational symmetric information bottleneck (DVSIB). DSIB, the deterministic limit of DVSIB, connects to modern contrastive learning approaches such as Barlow Twins, among others. We evaluate these methods on Noisy MNIST and Noisy CIFAR-100, showing that algorithms better matched to the structure of the problem like DVSIB and beta-DVCCA produce better latent spaces as measured by classification accuracy, dimensionality of the latent variables, sample efficiency, and consistently outperform other approaches under comparable conditions. Additionally, we benchmark against state-of-the-art models, achieving superior or competitive accuracy. Our results demonstrate that this framework can seamlessly incorporate diverse multi-view representation learning algorithms, providing a foundation for designing novel, problem-specific loss functions.
</description>
</item>

<item>
<title>
Diffeomorphism-based feature learning using Poincaré inequalities on augmented input space
</title>
<link>
http://jmlr.org/papers/v26/23-1707.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1707/23-1707.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Romain Verdière, Clémentine Prieur, Olivier Zahm</author>
<description>
We propose a gradient-enhanced algorithm for high-dimensional function approximation.
The algorithm proceeds in two steps: firstly, we reduce the input dimension by learning the relevant input features from gradient evaluations, and secondly, we regress the function output against the pre-learned features. To ensure theoretical guarantees, we construct the feature map as the first components of a diffeomorphism, which we learn by minimizing an error bound obtained using Poincaré Inequality applied either in the input space or in the feature space. This leads to two different strategies, which we compare both theoretically and numerically and relate to existing methods in the literature.
In addition, we propose a dimension augmentation trick to increase the approximation power of feature detection.
A generalization to vector-valued functions demonstrate that our methodology directly applies to learning autoencoders. Here, we approximate the identity function over a given dataset by a composition of feature map (encoder) with the regression function (decoder). In practice, we construct the diffeomorphism using coupling flows, a particular class of invertible neural networks.
Numerical experiments on various high-dimensional functions show that the proposed algorithm outperforms state-of-the-art competitors, especially with small datasets.
</description>
</item>

<item>
<title>
Finite Expression Method for Solving High-Dimensional Partial Differential Equations
</title>
<link>
http://jmlr.org/papers/v26/23-1290.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1290/23-1290.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Senwei Liang, Haizhao Yang</author>
<description>
Designing efficient and accurate numerical solvers for high-dimensional partial differential equations (PDEs) remains a challenging and important topic in computational science and engineering, mainly due to the &#34;curse of dimensionality&#34; in designing numerical schemes that scale in dimension. This paper introduces a new methodology that seeks an approximate PDE solution in the space of functions with finitely many analytic expressions and, hence, this methodology is named the finite expression method (FEX). It is proved in approximation theory that FEX can avoid the curse of dimensionality. As a proof of concept, a deep reinforcement learning method is proposed to implement FEX for various high-dimensional PDEs in different dimensions, achieving high and even machine accuracy with a memory complexity polynomial in dimension and an amenable time complexity. An approximate solution with finite analytic expressions also provides interpretable insights into the ground truth PDE solution, which can further help to advance the understanding of physical systems and design postprocessing techniques for a refined solution.
</description>
</item>

<item>
<title>
Randomly Projected Convex Clustering Model: Motivation, Realization, and Cluster Recovery Guarantees
</title>
<link>
http://jmlr.org/papers/v26/23-0384.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0384/23-0384.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Ziwen Wang, Yancheng Yuan, Jiaming Ma, Tieyong Zeng, Defeng Sun</author>
<description>
In this paper, we propose a randomly projected convex clustering model for clustering a collection of $n$ high dimensional data points in $\mathbb{R}^d$ with $K$ hidden clusters. Compared to the convex clustering model for clustering original data with dimension $d$, we prove that, under some mild conditions, the perfect recovery of the cluster membership assignments of the convex clustering model, if exists, can be preserved by the randomly projected convex clustering model with embedding dimension $m = O(\epsilon^{-2}\log(n))$, where $\epsilon &gt; 0$ is some given parameter. We further prove that the embedding dimension can be improved to be $O(\epsilon^{-2}\log(K))$, which is independent of the number of data points. We also establish the recovery guarantees of our proposed model with uniform weights for clustering a mixture of spherical Gaussians. Extensive numerical results demonstrate the robustness and superior performance of the randomly projected convex clustering model. The numerical results will also demonstrate that the randomly projected convex clustering model can outperform other popular clustering models on the dimension-reduced data, including the randomly projected K-means model.
</description>
</item>

<item>
<title>
Minimax Optimal Deep Neural Network Classifiers Under Smooth Decision Boundary
</title>
<link>
http://jmlr.org/papers/v26/22-0758.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0758/22-0758.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Tianyang Hu, Ruiqi Liu, Zuofeng Shang, Guang Cheng</author>
<description>
Deep learning has gained huge empirical successes in large-scale classification problems. In contrast, there is a lack of statistical understanding about deep learning methods, particularly in the minimax optimality perspective. For instance, in the classical smooth decision boundary setting, existing deep neural network (DNN) approaches are rate-suboptimal, and it remains elusive how to construct minimax optimal DNN classifiers. Moreover, it is interesting to explore whether DNN classifiers can circumvent the &#34;curse of dimensionality&#34; in handling high-dimensional data. The contributions of this paper are two-fold. First, based on a localized margin framework, we discover the source of suboptimality of existing DNN approaches. Motivated by this, we propose a new deep learning classifier using a divide-and-conquer technique: DNN classifiers are constructed on each local region and then aggregated to a global one. We further propose a localized version of the classical Tsybakov’s noise condition, under which statistical optimality of our new classifier is established. Second, we show that DNN classifiers can adapt to low-dimensional data structures and circumvent the “curse of dimensionality” in the sense that the minimax rate only depends on the effective dimension, potentially much smaller than the actual data dimension. Numerical experiments are conducted on simulated data to corroborate our theoretical results.
</description>
</item>

<item>
<title>
Optimal and Efficient Algorithms for Decentralized Online Convex Optimization
</title>
<link>
http://jmlr.org/papers/v26/24-2137.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-2137/24-2137.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yuanyu Wan, Tong Wei, Bo Xue, Mingli Song, Lijun Zhang</author>
<description>
We investigate decentralized online convex optimization (D-OCO), in which a set of local learners are required to minimize a sequence of global loss functions using only local computations and communications. Previous studies have established $O(n^{5/4}\rho^{-1/2}\sqrt{T})$ and ${O}(n^{3/2}\rho^{-1}\log T)$ regret bounds for convex and strongly convex functions respectively, where $n$ is the number of local learners, $\rho&lt;1$ is the spectral gap of the communication matrix, and $T$ is the time horizon. However, there exist large gaps from the existing lower bounds, i.e., $\Omega(n\sqrt{T})$ for convex functions and $\Omega(n)$ for strongly convex functions. To fill these gaps, in this paper, we first develop a novel D-OCO algorithm that can respectively reduce the regret bounds for convex and strongly convex functions to $\tilde{O}(n\rho^{-1/4}\sqrt{T})$ and $\tilde{O}(n\rho^{-1/2}\log T)$. The primary technique is to design an online accelerated gossip strategy that enjoys a faster average consensus among local learners. Furthermore, by carefully exploiting spectral properties of a specific network topology, we enhance the lower bounds for convex and strongly convex functions to $\Omega(n\rho^{-1/4}\sqrt{T})$ and $\Omega(n\rho^{-1/2}\log T)$, respectively. These results suggest that the regret of our algorithm is nearly optimal in terms of $T$, $n$, and $\rho$ for both convex and strongly convex functions. Finally, we propose a projection-free variant of our algorithm to efficiently handle practical applications with complex constraints. Our analysis reveals that the projection-free variant can achieve ${O}(nT^{3/4})$ and ${O}(nT^{2/3}(\log T)^{1/3})$ regret bounds for convex and strongly convex functions with nearly optimal $\tilde{O}(\rho^{-1/2}\sqrt{T})$ and $\tilde{O}(\rho^{-1/2}T^{1/3}(\log T)^{2/3})$ communication rounds, respectively.
</description>
</item>

<item>
<title>
Characterizing Dynamical Stability of Stochastic Gradient Descent in Overparameterized Learning
</title>
<link>
http://jmlr.org/papers/v26/24-1547.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1547/24-1547.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Dennis Chemnitz, Maximilian Engel</author>
<description>
For overparameterized optimization tasks, such as those found in modern machine learning, global minima are generally not unique. In order to understand generalization in these settings, it is vital to study to which minimum an optimization algorithm converges. The possibility of having minima that are unstable under the dynamics imposed by the optimization algorithm limits the potential minima that the algorithm can find. In this paper, we characterize the global minima that are dynamically stable/unstable for both deterministic and stochastic gradient descent (SGD). In particular, we introduce a characteristic Lyapunov exponent that depends on the local dynamics around a global minimum and rigorously prove that the sign of this Lyapunov exponent determines whether SGD can accumulate at the respective global minimum.
</description>
</item>

<item>
<title>
PREMAP: A Unifying PREiMage APproximation Framework for Neural Networks
</title>
<link>
http://jmlr.org/papers/v26/24-1297.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1297/24-1297.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Xiyue Zhang, Benjie Wang, Marta Kwiatkowska, Huan Zhang</author>
<description>
Most methods for neural network verification focus on bounding the image, i.e., set of outputs for a given input set. This can be used to, for example, check the robustness of neural network predictions to bounded perturbations of an input. However, verifying properties concerning the preimage, i.e., the set of inputs satisfying an output property, requires abstractions in the input space. We present a general framework for preimage abstraction that produces under- and over-approximations of any polyhedral output set. Our framework employs cheap parameterised linear relaxations of the neural network, together with an anytime refinement procedure that iteratively partitions the input region by splitting on input features and neurons. The effectiveness of our approach relies on carefully designed heuristics and optimisation objectives to achieve rapid improvements in the approximation volume. We evaluate our method on a range of tasks, demonstrating significant improvement in efficiency and scalability to high-input-dimensional image classification tasks compared to state-of-the-art techniques. Further, we showcase the application to quantitative verification and robustness analysis, presenting a sound and complete algorithm for the former and providing sound quantitative results for the latter.
</description>
</item>

<item>
<title>
Score-Aware Policy-Gradient and Performance Guarantees using Local Lyapunov Stability
</title>
<link>
http://jmlr.org/papers/v26/24-1009.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1009/24-1009.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Céline Comte, Matthieu Jonckheere, Jaron Sanders, Albert Senen-Cerda</author>
<description>
In this paper, we introduce a policy-gradient method for model-based reinforcement learning (RL) that exploits a type of stationary distributions commonly obtained from Markov decision processes (MDPs) in stochastic networks, queueing systems, and statistical mechanics. Specifically, when the stationary distribution of the MDP belongs to an exponential family that is parametrized by policy parameters, we can improve existing policy gradient methods for average-reward RL. Our key identification is a family of gradient estimators, called score-aware gradient estimators (SAGEs), that enable policy gradient estimation without relying on value-function estimation in the aforementioned setting. We show that SAGE-based policy-gradient locally converges, and we obtain its regret. This includes cases when the state space of the MDP is countable and unstable policies can exist. Under appropriate assumptions such as starting sufficiently close to a maximizer and the existence of a local Lyapunov function, the policy under SAGE-based stochastic gradient ascent has an overwhelming probability of converging to the associated optimal policy. Furthermore, we conduct a numerical comparison between a SAGE-based policy-gradient method and an actor-critic method on several examples inspired from stochastic networks, queueing systems, and models derived from statistical physics. Our results demonstrate that a SAGE-based method finds close-to-optimal policies faster than an actor-critic method.
</description>
</item>

<item>
<title>
On the O(sqrt(d)/T^(1/4)) Convergence Rate of RMSProp and Its Momentum Extension Measured by l_1 Norm
</title>
<link>
http://jmlr.org/papers/v26/24-0523.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0523/24-0523.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Huan Li, Yiming Dong, Zhouchen Lin</author>
<description>
Although adaptive gradient methods have been extensively used in deep learning, their convergence rates proved in the literature are all slower than that of SGD, particularly with respect to their dependence on the dimension. This paper considers the classical RMSProp and its momentum extension and establishes the convergence rate of $\frac{1}{T}\sum_{k=1}^TE\left[||\nabla f(\mathbf{x}^k)||_1\right]\leq O(\frac{\sqrt{d}C}{T^{1/4}})$ measured by $\ell_1$ norm without the bounded gradient assumption, where $d$ is the dimension of the optimization variable, $T$ is the iteration number, and $C$ is a constant identical to that appeared in the optimal convergence rate of SGD. Our convergence rate matches the lower bound with respect to all the coefficients except the dimension $d$. Since $||\mathbf{x}||_2\ll ||\mathbf{x}||_1\leq\sqrt{d}||\mathbf{x}||_2$ for problems with extremely large $d$, our convergence rate can be considered to be analogous to the $\frac{1}{T}\sum_{k=1}^TE\left[||\nabla f(\mathbf{x}^k)||_2\right]\leq O(\frac{C}{T^{1/4}})$ rate of SGD in the ideal case of $||\nabla f(\mathbf{x})||_1=\varTheta(\sqrt{d})||\nabla f(\mathbf{x})||_2$.
</description>
</item>

<item>
<title>
Categorical Semantics of Compositional Reinforcement Learning
</title>
<link>
http://jmlr.org/papers/v26/24-0197.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0197/24-0197.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Georgios Bakirtzis, Michail Savvas, Ufuk Topcu</author>
<description>
Compositional knowledge representations in reinforcement learning (RL) facilitate modular, interpretable, and safe task specifications. However, generating compositional models requires the characterization of minimal assumptions for the robustness of the compositionality feature, especially in the case of functional decompositions. Using a categorical point of view, we develop a knowledge representation framework for a compositional theory of RL. Our approach relies on the theoretical study of the category $\mathsf{MDP}$, whose objects are Markov decision processes (MDPs) acting as models of tasks. The categorical semantics models the compositionality of tasks through the application of pushout operations akin to combining puzzle pieces. As a practical application of these pushout operations, we introduce zig-zag diagrams that rely on the compositional guarantees engendered by the category $\mathsf{MDP}$. We further prove that properties of the category $\mathsf{MDP}$ unify concepts, such as enforcing safety requirements and exploiting symmetries, generalizing previous abstraction theories for RL.
</description>
</item>

<item>
<title>
Transformers from Diffusion: A Unified Framework for Neural Message Passing
</title>
<link>
http://jmlr.org/papers/v26/23-1672.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1672/23-1672.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Qitian Wu, David Wipf, Junchi Yan</author>
<description>
Learning representations for structured data with certain geometries (e.g., observed or unobserved) is a fundamental challenge, wherein message passing neural networks (MPNNs) have become a de facto class of model solutions. In this paper, inspired by physical systems, we propose an energy-constrained diffusion model, which integrates the inductive bias of diffusion on manifolds with layer-wise constraints of energy minimization. We identify that the diffusion operators have a one-to-one correspondence with the energy functions implicitly descended by the diffusion process, and the finite-difference iteration for solving the energy-constrained diffusion system induces the propagation layers of various types of MPNNs operating on observed or latent structures. This leads to a unified mathematical framework for common neural architectures whose computational flows can be cast as message passing (or its special case), including MLPs, GNNs, and Transformers. Building on these insights, we devise a new class of neural message passing models, dubbed diffusion-inspired Transformers (DIFFormer), whose global attention layers are derived from the principled energy-constrained diffusion framework. Across diverse datasets ranging from real-world networks to images, texts, and physical particles, we demonstrate that the new model achieves promising performance in scenarios where the data structures are observed (as a graph), partially observed, or entirely unobserved.
</description>
</item>

<item>
<title>
Optimal Sample Selection Through Uncertainty Estimation and Its Application in Deep Learning
</title>
<link>
http://jmlr.org/papers/v26/23-1160.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1160/23-1160.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yong Lin, Chen Liu, Chenlu Ye, Qing Lian, Yuan Yao, Tong Zhang</author>
<description>
Modern deep learning heavily relies on large labeled datasets, which often comse with high costs in terms of both manual labeling and computational resources. To mitigate these challenges, researchers have explored the use of informative subset selection techniques. In this study, we present a theoretically optimal solution for addressing both sampling with and without labels within the context of linear softmax regression. Our proposed method, COPS (unCertainty based OPtimal Sub-sampling), is designed to minimize the expected loss of a model trained on subsampled data. Unlike existing approaches that rely on explicit calculations of the inverse covariance matrix, which are not easily applicable to deep learning scenarios, COPS leverages the model&#39;s logits to estimate the sampling ratio. This sampling ratio is closely associated with model uncertainty and can be effectively applied to deep learning tasks. Furthermore, we address the challenge of model sensitivity to misspecification by incorporating a down-weighting approach for low-density samples, drawing inspiration from previous works. To assess the effectiveness of our proposed method, we conducted extensive empirical experiments using deep neural networks on benchmark datasets. The results consistently showcase the superior performance of COPS compared to baseline methods, reaffirming its efficacy.
</description>
</item>

<item>
<title>
Actor-Critic learning  for mean-field control in continuous time
</title>
<link>
http://jmlr.org/papers/v26/23-0345.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0345/23-0345.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Noufel FRIKHA, Maximilien GERMAIN, Mathieu LAURIERE, Huyen PHAM, Xuanye SONG</author>
<description>
We study policy gradient for mean-field control in continuous time in a  reinforcement learning setting. By considering randomised policies with entropy regularisation, we derive a gradient expectation representation of the value function, which is amenable to actor-critic type  algorithms, where the value functions and the policies are learnt alternately based on observation samples of the state  and model-free estimation of the population state distribution, either by offline or online learning. In the linear-quadratic mean-field framework, we obtain an exact parametrisation of the actor and critic functions defined on the Wasserstein space. Finally, we illustrate the results of our algorithms with some numerical experiments on  concrete examples.
</description>
</item>

<item>
<title>
Modelling Populations of Interaction Networks via Distance Metrics
</title>
<link>
http://jmlr.org/papers/v26/22-0706.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0706/22-0706.pdf
</pdf>
<pubDate>2025</pubDate>
<author>George Bolt, Simón Lunagómez, Christopher Nemeth</author>
<description>
Network data arises through the observation of relational information between a collection of entities, for example, friendships (relations) amongst a sample of people (entities). Traditionally, statistical models of such data have been developed to analyse a single network, that is, a single collection of entities and relations. More recently, attention has shifted to analysing samples of networks. A driving force has been the analysis of connectome data, arising in neuroscience applications, where a single network is observed for each patient in a study. These models typically assume, within each network, the entities are the units of observation, that is, more data equates to including more entities. However, an alternative paradigm considers relations—such as edges or paths—as the observational units, exemplified by email exchanges or user navigations across a website. This interaction network framework has generally been applied to single networks, without extending to the case where multiple such networks are observed, for instance, analysing navigation patterns from many users. Motivated by this gap, we propose a new Bayesian modelling framework to analyse such data. Our approach is based on practitioner-specified distance metrics between networks, allowing us to parameterise models analogous to Gaussian distributions in network space, using location and scale parameters. We address the key challenge of defining meaningful distances between interaction networks, proposing two new metrics with theoretical guarantees and practical computation strategies. To enable efficient Bayesian inference, we develop specialised Markov chain Monte Carlo (MCMC) algorithms within the involutive MCMC (iMCMC) framework, tailored to the doubly-intractable and discrete nature of the induced posteriors. Through simulation studies, we demonstrate the robustness and efficiency of our approach, and we showcase its applicability with a case study on a location-based social network (LSBN) dataset.
</description>
</item>

<item>
<title>
BitNet: 1-bit Pre-training for Large Language Models
</title>
<link>
http://jmlr.org/papers/v26/24-2050.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-2050/24-2050.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Hongyu Wang, Shuming Ma, Lingxiao Ma, Lei Wang, Wenhui Wang, Li Dong, Shaohan Huang, Huaijie Wang, Jilong Xue, Ruiping Wang, Yi Wu, Furu Wei</author>
<description>
The increasing size of large language models (LLMs) has posed challenges for deployment and raised concerns about environmental impact due to high energy consumption. Previous research typically applies quantization after pre-training. While these methods avoid the need for model retraining, they often cause notable accuracy loss at extremely low bit-widths. In this work, we explore the feasibility and scalability of 1-bit pre-training. We introduce BitNet b1 and BitNet b1.58, the scalable and stable 1-bit Transformer architecture designed for LLMs. Specifically, we introduce BitLinear as a drop-in replacement of the nn.Linear layer in order to train 1-bit weights from scratch. Experimental results show that BitNet b1 achieves competitive performance, compared to state-of-the-art 8-bit quantization methods and FP16 Transformer baselines. With the ternary weight, BitNet b1.58 matches the half-precision Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in terms of latency, memory, throughput, and energy consumption. More profoundly, BitNet defines a new scaling law and recipe for training new generations of LLMs that are both high-performance and cost-effective. It enables a new computation paradigm and opens the door for designing specific hardware optimized for 1-bit LLMs.
</description>
</item>

<item>
<title>
Physics-informed Kernel Learning
</title>
<link>
http://jmlr.org/papers/v26/24-1536.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1536/24-1536.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Nathan Doumèche, Francis Bach, Gérard Biau, Claire Boyer</author>
<description>
Physics-informed machine learning typically integrates physical priors into the learning process by minimizing a loss function that includes both a data-driven term and a partial differential equation (PDE) regularization. Building on the formulation of the problem as a kernel regression task, we use Fourier methods to approximate the associated kernel, and propose a tractable estimator that minimizes the physics-informed risk function. We refer to this approach as physics-informed kernel learning (PIKL). This framework provides theoretical guarantees, enabling the quantification of the physical prior’s impact on convergence speed. We demonstrate the numerical performance of the PIKL estimator through simulations, both in the context of hybrid modeling and in solving PDEs. In particular, we show that PIKL can outperform physics-informed neural networks in terms of both accuracy and computation time. Additionally, we identify cases where PIKL surpasses traditional PDE solvers, particularly in scenarios with noisy boundary conditions.
</description>
</item>

<item>
<title>
Last-iterate Convergence of Shuffling Momentum Gradient Method under the Kurdyka-Lojasiewicz Inequality
</title>
<link>
http://jmlr.org/papers/v26/24-1243.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1243/24-1243.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yuqing Liang, Dongpo Xu</author>
<description>
Shuffling gradient algorithms are extensively used to solve finite-sum optimization problems in machine learning. However, their theoretical properties still need to be further explored, especially the last-iterate convergence in the non-convex setting. In this paper, we study the last-iterate convergence behavior of shuffling momentum gradient (SMG) method, a shuffling gradient algorithm with momentum. Specifically, we focus on the non-convex scenario and provide theoretical guarantees under arbitrary shuffling strategies. For non-convex objectives, we achieve the convergence of gradient norms at the last-iterate, showing that every accumulation point of the iterative sequence is a stationary point of the non-convex problem. Our analysis also reveals that the function values of the last-iterate converge to a finite value. Additionally, we obtain the asymptotic convergence rates of gradient norms at the minimum-iterate. By employing a uniform without-replacement sampling strategy, we further achieve an improved convergence rate for the minimum-iterate output. Under the Kurdyka-Lojasiewicz (KL) inequality, we establish the challenging strong limit-point convergence results. In particular, we prove that the whole sequence of iterates exhibits convergence to a stationary point of the finite-sum problem. By choosing an appropriate stepsize, we also obtain the corresponding rate of last-iterate convergence, matching available results in the strongly convex setting. Given that the last iteration is typically preferred as the output of the algorithm in applied scenarios, this paper contributes to narrowing the gap between theory and practice.
</description>
</item>

<item>
<title>
Posterior and Variational Inference for Deep Neural Networks with Heavy-Tailed Weights
</title>
<link>
http://jmlr.org/papers/v26/24-0894.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0894/24-0894.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Paul Egels, Ismaël Castillo</author>
<description>
We consider deep neural networks in a Bayesian framework with a prior distribution sampling the network weights  at random. Following a recent idea of Agapiou and Castillo (2024), who show that heavy-tailed prior distributions achieve automatic adaptation to smoothness, we introduce a simple Bayesian deep learning prior based on heavy-tailed weights and ReLU activation. We show that the corresponding posterior distribution achieves near-optimal minimax contraction rates, simultaneously adaptive to both intrinsic dimension and smoothness of the underlying function, in a variety of contexts including nonparametric regression, geometric data and Besov spaces. While most works so far need a form of model selection built-in within the prior distribution, a key aspect of our approach is that it does not require to sample hyperparameters to learn the architecture of the network. We also provide variational Bayes counterparts of the results, that show that mean-field variational approximations still benefit from near-optimal theoretical support.
</description>
</item>

<item>
<title>
Maximum Causal Entropy IRL in Mean-Field Games and GNEP Framework for Forward RL
</title>
<link>
http://jmlr.org/papers/v26/24-0458.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0458/24-0458.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Berkay Anahtarci, Can Deha Kariksiz, Naci Saldi</author>
<description>
This paper explores the use of Maximum Causal Entropy Inverse Reinforcement Learning (IRL) within the context of discrete-time stationary Mean-Field Games (MFGs) characterized by finite state spaces and an infinite-horizon, discounted-reward setting. Although the resulting optimization problem is non-convex with respect to policies, we reformulate it as a convex optimization problem in terms of state-action occupation measures by leveraging the linear programming framework of Markov Decision Processes. Based on this convex reformulation, we introduce a gradient descent algorithm with a guaranteed convergence rate to efficiently compute the optimal solution. Moreover, we develop a new method that conceptualizes the MFG problem as a Generalized Nash Equilibrium Problem (GNEP), enabling effective computation of the mean-field equilibrium for forward reinforcement learning (RL) problems and marking an advancement in MFG solution techniques. We further illustrate the practical applicability of our GNEP approach by employing this algorithm to generate data for numerical MFG examples.
</description>
</item>

<item>
<title>
Degree of Interference: A General Framework For Causal Inference Under Interference
</title>
<link>
http://jmlr.org/papers/v26/24-0119.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0119/24-0119.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yuki Ohnishi, Bikram Karmakar, Arman Sabbaghi</author>
<description>
One core assumption typically adopted for valid causal inference is that of no interference between experimental units, i.e., the outcome of an experimental unit is unaffected by the treatments assigned to other experimental units. This assumption can be violated in real-life experiments, which significantly complicates the task of causal inference. As the number of potential outcomes increases, it becomes challenging to disentangle direct treatment effects from “spillover” effects. Current methodologies are lacking, as they cannot handle arbitrary, unknown interference structures to permit inference on causal estimands. We present a general framework to address the limitations of existing approaches. Our framework is based on the new concept of the “degree of interference” (DoI). The DoI is a unit-level latent variable that captures the latent structure of interference. We also develop a data augmentation algorithm that adopts a blocked Gibbs sampler and Bayesian nonparametric methodology to perform inferences on the estimands under our framework. We illustrate the DoI concept and properties of our Bayesian methodology via extensive simulation studies and an analysis of a randomized experiment investigating the impact of a cash transfer program for which interference is a critical concern. Ultimately, our framework enables us to infer causal effects without strong structural assumptions on interference.
</description>
</item>

<item>
<title>
Quantifying the Effectiveness of Linear Preconditioning in Markov Chain Monte Carlo
</title>
<link>
http://jmlr.org/papers/v26/23-1633.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1633/23-1633.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Max Hird, Samuel Livingstone</author>
<description>
We study linear preconditioning in Markov chain Monte Carlo. We consider the class of well-conditioned distributions, for which several mixing time bounds depend on the condition number  $\kappa$.  First we show that well-conditioned distributions exist for which $\kappa$ can be arbitrarily large and yet no linear preconditioner can reduce it.  We then impose two sets of extra assumptions under which a linear preconditioner can significantly reduce $\kappa$.  For the random walk Metropolis we further provide upper and lower bounds on the spectral gap with tight $1/\kappa$ dependence.  This allows us to give conditions under which linear preconditioning can provably increase the gap.  We then study popular preconditioners such as the covariance, its diagonal approximation, the Hessian at the mode, and the QR decomposition.  We show conditions under which each of these reduce $\kappa$ to near its minimum. We also show that the diagonal approach can in fact increase the condition number.  This is of interest as diagonal preconditioning is the default choice in well-known software packages.  We conclude with a numerical study comparing preconditioners in different models, and we show how proper preconditioning can greatly reduce compute time in Hamiltonian Monte Carlo.
</description>
</item>

<item>
<title>
Sparse SVM with Hard-Margin Loss: a Newton-Augmented Lagrangian Method in Reduced Dimensions
</title>
<link>
http://jmlr.org/papers/v26/23-0953.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0953/23-0953.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Penghe Zhang, Naihua Xiu, Hou-Duo Qi</author>
<description>
The hard-margin loss function has been at the core of the support vector machine research from the very beginning due to its generalization capability. On the other hand, the cardinality constraint has been widely used for feature selection, leading to sparse solutions. This paper studies the sparse SVM with the hard-margin loss that integrates the virtues of both worlds, resulting in one of the most challenging models to solve. We cast the problem as a composite optimization with the cardinality constraint. We characterize its local minimizers in terms of pseudo KKT point that well captures the combinatorial structure of the problem, and investigate a sharper P-stationary point with a concise representation for algorithm design. We further develop an inexact proximal augmented Lagrangian method (iPAL). The different parts of the inexactness measurements from the {\rm P}-stationarity are controlled at different scales in a way that the generated sequence converges both globally and at a linear rate. To make iPAL practically efficient, we propose a gradient-Newton method in a subspace for the iPAL subproblem. This is accomplished by detecting active samples and features with the help of the proximal operator of the hard margin loss and the projection of the cardinality constraint. Extensive numerical results on both simulated and real data sets demonstrate that the proposed method is fast, produces sparse solution of high accuracy, and can lead to effective reduction on active samples and features  when compared with several leading solvers.
</description>
</item>

<item>
<title>
On Model Identification and Out-of-Sample Prediction of PCR with Applications to Synthetic Controls
</title>
<link>
http://jmlr.org/papers/v26/23-0102.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0102/23-0102.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Anish Agarwal, Devavrat Shah, Dennis Shen</author>
<description>
We analyze principal component regression (PCR) in a high-dimensional error-in-variables setting with fixed design. Under suitable conditions, we show that PCR consistently identifies the unique model with minimum $\ell_2$-norm. These results enable us to establish non-asymptotic out-of-sample prediction guarantees that improve upon the best known rates. In the course of our analysis, we introduce a natural linear algebraic condition between the in- and out-of-sample covariates, which allows us to avoid distributional assumptions for out-of-sample predictions. Our simulations illustrate the importance of this condition for generalization, even under covariate shifts. Accordingly, we construct a hypothesis test to check when this condition holds in practice. As a byproduct, our results also lead to novel results for the synthetic controls literature, a leading approach for policy evaluation. To the best of our knowledge, our prediction guarantees for the fixed design setting have been elusive in both the high-dimensional error-in-variables and synthetic controls literatures.
</description>
</item>

<item>
<title>
Bayesian Scalar-on-Image Regression with a Spatially Varying Single-layer Neural Network Prior
</title>
<link>
http://jmlr.org/papers/v26/22-0246.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0246/22-0246.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Ben Wu, Keru Wu, Jian Kang</author>
<description>
Deep neural networks (DNN) have been widely used in scalar-on-image regression to predict an outcome variable from imaging predictors.  However, training DNN typically requires large sample sizes for accurate prediction, and the resulting models often lack interpretability. In this work, we propose a novel Bayesian nonlinear scalar-on-image regression framework with a spatially varying single-layer neural network (SV-NN) prior. The SV-NN is constructed using a single hidden layer neural network with its weights generated by the soft-thresholded Gaussian process. Our framework enables the selection of interpretable image regions while achieving high prediction accuracy with limited training samples. The SV-NN offers large prior support for the imaging effect function, facilitating efficient posterior inference on image region selection and automatic network structures determination.  We establish the posterior consistency for model parameters and selection consistency for image regions when the number of voxels/pixels grows much faster than the sample size.  To ensure computational efficiency, we develop a stochastic gradient Langevin dynamics (SGLD) algorithm for posterior inference. We evaluate our method through extensive comparisons with state-of-the-art deep learning approaches, analyzing multiple real datasets, including task fMRI data from the Adolescent Brain Cognitive Development (ABCD) study.
</description>
</item>

<item>
<title>
DRM Revisited: A Complete Error Analysis
</title>
<link>
http://jmlr.org/papers/v26/24-1258.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1258/24-1258.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yuling Jiao, Ruoxuan Li, Peiying Wu, Jerry Zhijian Yang, Pingwen Zhang</author>
<description>
It is widely known that the error analysis for deep learning involves approximation, statistical, and optimization errors. However, it is challenging to combine them together due to overparameterization. In this paper, we address this gap by providing a comprehensive error analysis of the Deep Ritz Method (DRM). Specifically, we investigate a foundational question in the theoretical analysis of DRM under the overparameterized regime: given a target precision level, how can one determine the appropriate number of training samples, the key architectural parameters of the neural networks, the step size for the projected gradient descent optimization procedure, and the requisite number of iterations, such that the output of the gradient descent process closely approximates the true solution of the underlying partial differential equation to the specified precision?
</description>
</item>

<item>
<title>
Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF
</title>
<link>
http://jmlr.org/papers/v26/24-0720.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0720/24-0720.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Han Shen, Zhuoran Yang, Tianyi Chen</author>
<description>
Bilevel optimization has been recently applied to many machine learning tasks. However, their applications have been restricted to the supervised learning setting, where static objective functions with benign structures are considered. But bilevel problems such as incentive design, inverse reinforcement learning (RL), and RL from human feedback (RLHF) are often modeled as dynamic objective functions that go beyond the simple static objective structures, which pose significant challenges of using existing bilevel solutions. To tackle this new class of bilevel problems, we introduce the first principled algorithmic framework for solving bilevel RL problems through the lens of penalty formulation. We provide theoretical studies of the problem landscape and its penalty-based (policy) gradient algorithms. We demonstrate the effectiveness of our algorithms via simulations in the Stackelberg Markov game, RL from human feedback and incentive design.
</description>
</item>

<item>
<title>
Precise High-Dimensional Asymptotics for Quantifying Heterogeneous Transfers
</title>
<link>
http://jmlr.org/papers/v26/24-0454.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0454/24-0454.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Fan Yang, Hongyang R. Zhang, Sen Wu, Christopher Re, Weijie J. Su</author>
<description>
The problem of learning one task using samples from another task is central to transfer learning. In this paper, we focus on answering the following question: when does combining the samples from two related tasks perform better than learning with one target task alone? This question is motivated by an empirical phenomenon known as negative transfer often observed in transfer learning practice. While the transfer effect from one task to another depends on factors such as their sample sizes and the spectrum of their covariance matrices, precisely quantifying this dependence has remained a challenging problem. In order to compare a transfer learning estimator to single-task learning, one needs to compare the risks between the two estimators precisely. Further, the comparison depends on the distribution shifts between the two tasks. This paper applies recent developments of random matrix theory to tackle this challenge in a high-dimensional linear regression setting with two tasks. We provide precise high-dimensional asymptotics for the bias and variance of a classical hard parameter sharing (HPS) estimator in the proportional limit, when the sample sizes of both tasks increase proportionally with dimension at fixed ratios. The precise asymptotics apply to various types of distribution shifts, including covariate shifts, model shifts, and combinations of both. We illustrate these results in a random-effects model to mathematically prove a phase transition from positive to negative transfer as the number of source task samples increases. One insight from the analysis is that a rebalanced HPS estimator, which downsizes the source task when the model shift is high, achieves the minimax optimal rate. The finding regarding phase transition also applies to multiple tasks when feature covariates are shared across all tasks. Simulations validate the accuracy of the high-dimensional asymptotics for finite dimensions.
</description>
</item>

<item>
<title>
Score-based Causal Representation Learning: Linear and General Transformations
</title>
<link>
http://jmlr.org/papers/v26/24-0194.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0194/24-0194.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Burak Varici, Emre Acartürk, Karthikeyan Shanmugam, Abhishek Kumar, Ali Tajer</author>
<description>
This paper addresses intervention-based causal representation learning (CRL) under a general nonparametric latent causal model and an unknown transformation that maps the latent variables to the observed variables. Linear and general transformations are investigated. The paper addresses both the identifiability and achievability aspects. Identifiability refers to determining algorithm-agnostic conditions that ensure the recovery of the true latent causal variables and the underlying latent causal graph. Achievability refers to the algorithmic aspects and addresses designing algorithms that achieve identifiability guarantees. By drawing novel connections between score functions (i.e., the gradients of the logarithm of density functions) and CRL, this paper designs a score-based class of algorithms that ensures both identifiability and achievability. First, the paper focuses on linear transformations and shows that one stochastic hard intervention per node suffices to guarantee identifiability. It also provides partial identifiability guarantees for soft interventions, including identifiability up to mixing with parents for general causal models and perfect recovery of the latent graph for sufficiently nonlinear causal models. Secondly, it focuses on general transformations and demonstrates that two stochastic hard interventions per node are sufficient for identifiability. This is achieved by defining a differentiable loss function whose global optima ensure identifiability for general CRL. Notably, one does not need to know which pair of interventional environments has the same node intervened. Finally, the theoretical results are empirically validated via experiments on structured synthetic data and image data.
</description>
</item>

<item>
<title>
On the Statistical Properties of Generative Adversarial Models for Low Intrinsic Data Dimension
</title>
<link>
http://jmlr.org/papers/v26/24-0054.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0054/24-0054.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Saptarshi Chakraborty, Peter L. Bartlett</author>
<description>
Despite the remarkable empirical successes of Generative Adversarial Networks (GANs), the theoretical guarantees for their statistical accuracy remain rather pessimistic. In particular, the data distributions on which GANs are applied, such as natural images, are often hypothesized to have an intrinsic low-dimensional structure in a typically high-dimensional feature space, but this is often not reflected in the derived rates in the state-of-the-art analyses. In this paper, we attempt to bridge the gap between the theory and practice of GANs and their bidirectional variant, Bi-directional GANs (BiGANs), by deriving statistical guarantees on the estimated densities in terms of the intrinsic dimension of the data and the latent space. We analytically show that if one has access to $n$ samples from the unknown target distribution and the network architectures are properly chosen, the expected Wasserstein-1 distance of the estimates from the target scales as $O\left( n^{-1/d_\mu } \right)$  for GANs and $\tilde{O}\left( n^{-1/(d_\mu+\ell)} \right)$  for BiGANs,  where $d_\mu$ and $\ell$ are the upper Wasserstein-1 dimension of the data-distribution and latent-space dimension, respectively. The theoretical analyses not only suggest that these methods successfully avoid the curse of dimensionality, in the sense that the exponent of $n$ in the error rates does not depend on the data dimension but also serve to bridge the gap between the theoretical analyses of GANs and the known sharp rates from optimal transport literature.  Additionally, we demonstrate that GANs can effectively achieve the minimax optimal rate even for non-smooth underlying distributions, with the use of interpolating generator networks.
</description>
</item>

<item>
<title>
Prominent Roles of Conditionally Invariant Components in Domain Adaptation: Theory and Algorithms
</title>
<link>
http://jmlr.org/papers/v26/23-1234.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1234/23-1234.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Keru Wu, Yuansi Chen, Wooseok Ha, Bin Yu</author>
<description>
Domain adaptation (DA) is a statistical learning problem that arises when the distribution of the source data used to train a model differs from that of the target data used to evaluate the model. While many DA algorithms have demonstrated considerable empirical success, blindly applying these algorithms can often lead to worse performance on new datasets. To address this, it is crucial to clarify the assumptions under which a DA algorithm has good target performance. In this work, we focus on the assumption of the presence of conditionally invariant components (CICs), which are relevant for prediction and remain conditionally invariant across the source and target data. We demonstrate that CICs, which can be estimated through conditional invariant penalty (CIP), play three prominent roles in providing target risk guarantees in DA.  First, we propose a new algorithm based on CICs, importance-weighted conditional invariant penalty (IW-CIP), which has target risk guarantees beyond simple settings such as covariate shift and label shift. Second, we show that CICs help identify large discrepancies between source and target risks of other DA algorithms. Finally, we demonstrate that incorporating CICs into the domain invariant projection (DIP) algorithm can address its failure scenario caused by label-flipping features. We support our new algorithms and theoretical findings via numerical experiments on synthetic data, MNIST, CelebA, Camelyon17, and DomainNet datasets.
</description>
</item>

<item>
<title>
Near-Optimal Nonconvex-Strongly-Convex Bilevel Optimization with Fully First-Order Oracles
</title>
<link>
http://jmlr.org/papers/v26/23-1104.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1104/23-1104.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Lesi Chen, Yaohua Ma, Jingzhao Zhang</author>
<description>
In this work, we consider bilevel optimization when the lower-level problem is strongly convex. Recent works show that with a Hessian-vector product (HVP) oracle, one can provably find an $\epsilon$-stationary point within ${O}(\epsilon^{-2})$ oracle calls. However, the HVP oracle may be inaccessible or expensive in practice. Kwon et al. (ICML 2023) addressed this issue by proposing a first-order method that can achieve the same goal at a slower rate of $\tilde{O}(\epsilon^{-3})$.  In this paper, we incorporate a two-time-scale update to improve their method to achieve the near-optimal $\tilde{O}(\epsilon^{-2})$ first-order oracle complexity. Our analysis is highly extensible. In the stochastic setting, our algorithm can achieve the stochastic first-order oracle complexity of $\tilde {O}(\epsilon^{-4})$ and $\tilde {O}(\epsilon^{-6})$ when the stochastic noises are only in the  upper-level objective and in both level objectives, respectively.  When the objectives have higher-order smoothness conditions, our deterministic method can escape saddle points by injecting noise, and can be accelerated to  achieve a faster rate of $\tilde {O}(\epsilon^{-1.75})$ using Nesterov&#39;s momentum.
</description>
</item>

<item>
<title>
Adaptive Distributed Kernel Ridge Regression: A Feasible Distributed Learning Scheme for Data Silos
</title>
<link>
http://jmlr.org/papers/v26/23-0806.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0806/23-0806.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Shao-Bo Lin, Xiaotong Liu, Di Wang, Hai Zhang, Ding-Xuan Zhou</author>
<description>
Data silos, mainly caused by privacy and interoperability, significantly constrain collaborations among different organizations with similar data for the same purpose. Distributed learning based on divide-and-conquer provides a promising way to settle the data silos, but it suffers from several challenges, including autonomy, privacy guarantees, and the necessity of collaborations. This paper focuses on developing an adaptive distributed kernel ridge regression (AdaDKRR) by taking autonomy in parameter selection, privacy in communicating non-sensitive information, and the necessity of collaborations for performance improvement into account. We provide both solid theoretical verifications and comprehensive experiments for AdaDKRR to demonstrate its feasibility and effectiveness. Theoretically, we prove that under some mild conditions, AdaDKRR performs similarly to running the optimal learning algorithms on the whole data, verifying the necessity of collaborations and showing that no other distributed learning scheme can essentially beat AdaDKRR under the same conditions. Numerically, we test AdaDKRR on both toy simulations and two real-world applications to show that AdaDKRR is superior to other existing distributed learning schemes. All these results show that AdaDKRR is a feasible scheme to overcome data silos, which are highly desired in numerous application regions such as intelligent decision-making, pricing forecasting, and performance prediction for products.
</description>
</item>

<item>
<title>
On Global and Local Convergence of Iterative Linear Quadratic Optimization Algorithms for Discrete Time Nonlinear Control
</title>
<link>
http://jmlr.org/papers/v26/22-1271.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-1271/22-1271.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Vincent Roulet, Siddhartha Srinivasa, Maryam Fazel, Zaid Harchaoui</author>
<description>
A classical approach for solving discrete time nonlinear control on a finite horizon consists in repeatedly minimizing linear quadratic approximations of the original problem around current candidate solutions. While widely popular in many domains, such an approach has mainly been analyzed locally. We provide detailed convergence guarantees to stationary points as well as local linear convergence rates for the Iterative Linear Quadratic Regulator (ILQR) algorithm and its Differential Dynamic Programming (DDP) variant. For problems without costs on control variables, we observe that global convergence to minima can be ensured provided that the linearized discrete time dynamics are surjective, costs on the state variables are gradient dominated. We further detail quadratic local convergence when the costs are self-concordant. We show that surjectivity of the linearized dynamics hold for appropriate discretization schemes given the existence of a feedback linearization scheme. We present complexity bounds of algorithms based on linear quadratic approximations through the lens of generalized Gauss-Newton methods. Our analysis uncovers several convergence phases for regularized generalized Gauss-Newton algorithms.
</description>
</item>

<item>
<title>
A Decentralized Proximal Gradient Tracking Algorithm for Composite Optimization on Riemannian Manifolds
</title>
<link>
http://jmlr.org/papers/v26/24-1989.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1989/24-1989.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Lei Wang, Le Bao, Xin Liu</author>
<description>
This paper focuses on minimizing a smooth function combined with a nonsmooth regularization term on a compact Riemannian submanifold embedded in the Euclidean space under a decentralized setting. Typically, there are two types of approaches at present for tackling such composite optimization problems. The first, subgradient-based approaches, rely on subgradient information of the objective function to update variables, achieving an iteration complexity of $O(\epsilon^{-4}\log^2(\epsilon^{-2}))$. The second, smoothing approaches, involve constructing a smooth approximation of the nonsmooth regularization term, resulting in an iteration complexity of $O(\epsilon^{-4})$. This paper proposes a proximal gradient type algorithm that fully exploits the composite structure. The global convergence to a stationary point is established with a significantly improved iteration complexity of $O(\epsilon^{-2})$. To validate the effectiveness and efficiency of our proposed method, we present numerical results from real-world applications, showcasing its superior performance compared to existing approaches.
</description>
</item>

<item>
<title>
Learning conditional distributions on continuous spaces
</title>
<link>
http://jmlr.org/papers/v26/24-0924.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0924/24-0924.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Cyril Benezet, Ziteng Cheng, Sebastian Jaimungal</author>
<description>
We investigate sample-based learning of conditional distributions on multi-dimensional unit boxes, allowing for different dimensions of the feature and target spaces. Our approach involves clustering data near varying query points in the feature space to create empirical measures in the target space. We employ two distinct clustering schemes: one based on a fixed-radius ball and the other on nearest neighbors. We establish upper bounds for the convergence rates of both methods and, from these bounds, deduce optimal configurations for the radius and the number of neighbors. We propose to incorporate the nearest neighbors method into neural network training, as our empirical analysis indicates it has better performance in practice. For efficiency, our training process utilizes approximate nearest neighbors search with random binary space partitioning. Additionally, we employ the Sinkhorn algorithm and a sparsity-enforced transport plan. Our empirical findings demonstrate that, with a suitably designed structure, the neural network has the ability to adapt to a suitable level of Lipschitz continuity locally.
</description>
</item>

<item>
<title>
A Unified Analysis of Nonstochastic Delayed Feedback for Combinatorial Semi-Bandits, Linear Bandits, and MDPs
</title>
<link>
http://jmlr.org/papers/v26/24-0496.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0496/24-0496.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Lukas Zierahn, Dirk van der Hoeven, Tal Lancewicki, Aviv Rosenberg, Nicolò Cesa-Bianchi</author>
<description>
We derive a new analysis of Follow The Regularized Leader (FTRL) for online learning with delayed bandit feedback. By separating the cost of delayed feedback from that of bandit feedback, our analysis allows us to obtain new results in four important settings. We derive the first optimal (up to logarithmic factors) regret bounds for combinatorial semi-bandits with delay and adversarial Markov Decision Processes with delay (both known and unknown transition functions). 
Furthermore, we use our analysis to develop an efficient algorithm for linear bandits with delay achieving near-optimal regret bounds. In order to derive these results we show that FTRL remains stable across multiple rounds under mild assumptions on the regularizer.
</description>
</item>

<item>
<title>
Error bounds for particle gradient descent, and extensions of the log-Sobolev and Talagrand inequalities
</title>
<link>
http://jmlr.org/papers/v26/24-0437.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0437/24-0437.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Rocco Caprio, Juan Kuntz, Samuel Power, Adam M. Johansen</author>
<description>
We derive non-asymptotic error bounds for particle gradient descent (PGD,  Kuntz et al. (2023)), a recently introduced algorithm for maximum likelihood estimation of large latent variable models obtained by discretizing a gradient flow of the free energy.  We begin by showing that the flow converges exponentially fast to the free energy&#39;s minimizers for models satisfying a condition that generalizes both the log-Sobolev and the Polyak--Łojasiewicz inequalities (LSI and PŁI, respectively). We achieve this by extending a result well-known in the optimal transport literature (that the LSI implies the Talagrand inequality) and its counterpart in the optimization literature (that the PŁI implies the so-called quadratic growth condition), and applying the extension to our new setting. We also generalize the Bakry--Émery Theorem and show that the LSI/PŁI  extension holds for models with strongly concave log-likelihoods. For such models, we further control PGD&#39;s discretization error and obtain the non-asymptotic error bounds. While we are motivated by the study of PGD, we believe that the inequalities and results we extend may be of independent interest.
</description>
</item>

<item>
<title>
Linear Hypothesis Testing in High-Dimensional Expected Shortfall Regression with Heavy-Tailed Errors
</title>
<link>
http://jmlr.org/papers/v26/24-0061.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0061/24-0061.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Gaoyu Wu, Jelena Bradic, Kean Ming Tan, Wen-Xin Zhou</author>
<description>
Expected shortfall (ES) is widely used for characterizing the tail of a distribution across various fields, particularly in financial risk management. In this paper, we explore a two-step procedure that leverages an orthogonality property to reduce sensitivity to nuisance parameters when estimating within a joint quantile and expected shortfall regression framework. For high-dimensional sparse models, we propose a robust $\ell_1$-penalized two-step approach capable of handling heavy-tailed data distributions. We establish non-asymptotic estimation error bounds and propose an appropriate growth rate for the diverging robustification parameter. To facilitate statistical inference for certain linear combinations of the ES regression coefficients, we construct debiased estimators and develop their asymptotic distributions, which form the basis for constructing valid confidence intervals. We validate the proposed method through simulation studies, demonstrating its effectiveness in high-dimensional linear models with heavy-tailed errors.
</description>
</item>

<item>
<title>
Efficient Numerical Integration in Reproducing Kernel Hilbert Spaces via Leverage Scores Sampling
</title>
<link>
http://jmlr.org/papers/v26/23-1551.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1551/23-1551.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Antoine Chatalic, Nicolas Schreuder, Ernesto De Vito, Lorenzo Rosasco</author>
<description>
In this work we consider the problem of numerical integration, i.e., approximating integrals with respect to a target probability measure using only pointwise evaluations of the integrand. We focus on the setting in which the target distribution is only accessible through a set of $n$ i.i.d. observations, and the integrand belongs to a reproducing kernel Hilbert space. We propose an efficient procedure which exploits a small i.i.d. random subset of $m \lt n$ samples drawn either uniformly or using approximate leverage scores from the initial observations. Our main result is an upper bound on the approximation error of this procedure for both sampling strategies. It yields sufficient conditions on the subsample size to recover the standard (optimal) $n^{-1/2}$ rate while reducing drastically the number of functions evaluations---and thus the overall computational cost. Moreover, we obtain rates with respect to the number $m$ of evaluations of the integrand which adapt to its smoothness, and match known optimal rates for instance for Sobolev spaces. We illustrate our theoretical findings with numerical experiments on real datasets, which highlight the attractive efficiency-accuracy tradeoff of our method compared to existing randomized and greedy quadrature methods. We note that, the problem of numerical integration in RKHS amounts to designing a discrete approximation of the kernel mean embedding of the target distribution. As a consequence, direct applications of our results also include the efficient computation of maximum mean discrepancies between distributions and the design of efficient kernel-based tests.
</description>
</item>

<item>
<title>
Distribution Free Tests for Model Selection Based on Maximum Mean Discrepancy with Estimated Parameters
</title>
<link>
http://jmlr.org/papers/v26/23-1199.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1199/23-1199.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Florian Brück, Jean-David Fermanian, Aleksey Min</author>
<description>
There exist several testing procedures based on the maximum mean discrepancy (MMD) to address the challenge of model specification. However, these testing procedures ignore the presence of estimated parameters in the case of composite null hypotheses. In this paper, we first illustrate the effect of parameter estimation in model specification tests based on the MMD. Second, we propose simple model specification and model selection tests in the case of models with estimated parameters. All our tests are asymptotically standard normal under the null, even when the true underlying distribution belongs to the competing parametric families. A simulation study and a real data analysis illustrate the performance of our tests in terms of power and level.
</description>
</item>

<item>
<title>
Statistical field theory for Markov decision processes under uncertainty
</title>
<link>
http://jmlr.org/papers/v26/23-0905.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0905/23-0905.pdf
</pdf>
<pubDate>2025</pubDate>
<author>George Stamatescu</author>
<description>
A statistical field theory is introduced for finite state and action Markov decision processes with unknown parameters, in a Bayesian setting. The Bellman equation, for policy evaluation and the optimal value function in finite and discounted infinite horizon problems, is studied as a disordered interacting dynamical system. The Markov decision process transition probabilities and mean-rewards are interpreted as quenched random variables and the value functions, or the iterates of the Bellman equation, are deterministic variables that evolve dynamically. The posterior over value functions is then equivalent to the quenched average of Fourier inverse of the Martin-Siggia-Rose-De Dominicis-Janssen generating function. The formalism enables the use of methods from field theory to compute posterior moments of value functions. The paper presents two such methods, corresponding to two distinct asymptotic limits. First, the classical approximation is applied, corresponding to the asymptotic data limit. This approximation recovers so-called plug-in estimators for the mean of the value functions. Second, a dynamic mean field theory is derived, showing that under certain assumptions the state-action values are statistically independent across state-action pairs in the asymptotic state space limit. The state-action value statistics can be computed from a set of self-consistent mean field equations, which we call dynamic mean field programming (DMFP). Collectively, the results provide analytic insight into the structure of model uncertainty in Markov decision processes, and pave the way toward more advanced field theoretic techniques and applications to planning and reinforcement learning problems.
</description>
</item>

<item>
<title>
Bayesian Data Sketching for Varying Coefficient Regression Models
</title>
<link>
http://jmlr.org/papers/v26/23-0505.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0505/23-0505.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Rajarshi Guhaniyogi, Laura Baracaldo, Sudipto Banerjee</author>
<description>
Varying coefficient models are popular for estimating nonlinear regression functions in functional data models. Their Bayesian variants have received limited attention in large data applications, primarily due to prohibitively slow posterior computations using Markov chain Monte Carlo (MCMC) algorithms. We introduce Bayesian data sketching for varying coefficient models to obviate computational challenges presented by large sample sizes. To address the challenges of analyzing large data, we compress the functional response vector and predictor matrix by a random linear transformation to achieve dimension reduction and conduct inference on the compressed data. Our approach distinguishes itself from several existing methods for analyzing large functional data in that it requires neither the development of new models or algorithms nor any specialized computational hardware while delivering fully model-based Bayesian inference. Well-established methods and algorithms for varying-coefficient regression models can be applied to the compressed data. We establish posterior contraction rates for estimating the varying coefficients and predicting the outcome at new locations with the randomly compressed data model. We use simulation experiments and analyze remote sensed vegetation data to empirically illustrate the inferential and computational efficiency of our approach.
</description>
</item>

<item>
<title>
Bagged k-Distance for Mode-Based Clustering  Using the Probability of Localized Level Sets
</title>
<link>
http://jmlr.org/papers/v26/22-1179.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-1179/22-1179.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Hanyuan Hang</author>
<description>
In this paper, we propose an ensemble learning algorithm named bagged $k$-distance for mode-based clustering (BDMBC) by putting forward a new measure called the probability of localized level sets (PLLS), which enables us to find all clusters for varying densities with a global threshold. On the theoretical side, we show that with a properly chosen number of nearest neighbors $k_D$ in the bagged $k$-distance, the sub-sample size $s$, the bagging rounds $B$, and the number of nearest neighbors $k_L$ for the localized level sets, BDMBC can achieve optimal convergence rates for mode estimation. It turns out that with a relatively small $B$, the sub-sample size $s$ can be much smaller than the number of training data $n$ at each bagging round, and the number of nearest neighbors $k_D$ can be reduced simultaneously. Moreover, we establish fast convergence rates for the level set estimation of the PLLS in terms of Hausdorff distance, which reveals that BDMBC can find localized level sets for varying densities and thus enjoys local adaptivity. On the practical side, we conduct numerical experiments to empirically verify the effectiveness of BDMBC for mode estimation and level set estimation, which demonstrates the promising accuracy and efficiency of our proposed algorithm.
</description>
</item>

<item>
<title>
Linear cost and exponentially convergent approximation of Gaussian Matérn processes on intervals
</title>
<link>
http://jmlr.org/papers/v26/24-1779.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1779/24-1779.pdf
</pdf>
<pubDate>2025</pubDate>
<author>David Bolin, Vaibhav Mehandiratta, Alexandre B. Simas</author>
<description>
The computational cost for inference and prediction of statistical models based on Gaussian processes with Matérn covariance functions scales cubically with the number of observations, limiting their applicability to large data sets. The cost can be reduced in certain special cases, but there are no generally applicable exact methods with linear cost. Several approximate methods have been introduced to reduce the cost, but most lack theoretical guarantees for accuracy. We consider Gaussian processes on bounded intervals with Matérn covariance functions and, for the first time, develop a generally applicable method with linear cost and a covariance error that decreases exponentially fast in the order $m$ of the proposed approximation. The method is based on an optimal rational approximation of the spectral density and results in an approximation that can be represented as a sum of $m$ independent Gaussian Markov processes, facilitating usage in general software for statistical inference. Besides theoretical justifications, we demonstrate accuracy empirically through carefully designed simulation studies, which show that the method outperforms state-of-the-art alternatives in accuracy for fixed computational cost in tasks like Gaussian process regression.
</description>
</item>

<item>
<title>
Invariant Subspace Decomposition
</title>
<link>
http://jmlr.org/papers/v26/24-0699.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0699/24-0699.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Margherita Lazzaretto, Jonas Peters, Niklas Pfister</author>
<description>
We consider the task of predicting a response $Y$ from a set of covariates $X$ in settings where the conditional distribution of $Y$ given $X$ changes over time. For this to be feasible, assumptions on how the conditional distribution changes over time are required. Existing approaches assume, for example, that changes occur smoothly over time so that short-term prediction using only the recent past becomes feasible. To additionally exploit observations further in the past, we propose a novel invariance-based framework for linear conditionals, called Invariant Subspace Decomposition (ISD), that splits the conditional distribution into a time-invariant and a residual time-dependent component. As we show, this decomposition can be employed both for zero-shot and time-adaptation prediction tasks, that is, settings where either no or a small amount of training data is available at the time points we want to predict $Y$ at, respectively. We propose a practical estimation procedure, which automatically infers the decomposition using tools from approximate joint matrix diagonalization. Furthermore, we provide finite sample guarantees for the proposed estimator and demonstrate empirically that it indeed improves on approaches that do not use the additional invariant structure.
</description>
</item>

<item>
<title>
Posterior Concentrations of Fully-Connected Bayesian Neural Networks with General Priors on the Weights
</title>
<link>
http://jmlr.org/papers/v26/24-0425.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0425/24-0425.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Insung Kong, Yongdai Kim</author>
<description>
Bayesian approaches for training deep neural networks (BNNs) have received significant interest and have been effectively utilized in a wide range of applications. Several studies have examined the properties of posterior concentrations in BNNs. However, most of these studies focus solely on BNN models with sparse or heavy-tailed priors. Surprisingly, there are currently no theoretical results for BNNs using Gaussian priors, which are the most commonly used in practice. The lack of theory arises from the absence of approximation results of Deep Neural Networks (DNNs) that are non-sparse and have bounded parameters. In this paper, we present a new approximation theory for non-sparse DNNs with bounded parameters. Additionally, based on the approximation theory,  we show that BNNs with non-sparse general priors can achieve near-minimax optimal posterior concentration rates around the true model.
</description>
</item>

<item>
<title>
Outlier Robust and Sparse Estimation of Linear Regression Coefficients
</title>
<link>
http://jmlr.org/papers/v26/23-1583.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1583/23-1583.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Takeyuki Sasai, Hironori Fujisawa</author>
<description>
We consider outlier-robust and sparse estimation of linear regression coefficients, when the covariates and the noises are contaminated by adversarial outliers and noises are sampled from a heavy-tailed distribution. Our results present sharper error bounds under weaker assumptions than prior studies that share similar interests with this study. Our analysis relies on some sharp concentration inequalities resulting from generic chaining.
</description>
</item>

<item>
<title>
Affine Rank Minimization via Asymptotic Log-Det Iteratively Reweighted Least Squares
</title>
<link>
http://jmlr.org/papers/v26/23-0943.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0943/23-0943.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Sebastian Krämer</author>
<description>
The affine rank minimization problem is a well-known approach to matrix recovery. While there are various surrogates to this NP-hard problem, we prove that the asymptotic minimization of log-det objective functions indeed always reveals the desired, lowest-rank matrices---whereas such may or may not recover a sought-after ground truth. Concerning commonly applied methods such as iteratively reweighted least squares, one thus remains with two difficult to distinguish concerns: how problematic are local minima inherent to the approach truly; and opposingly, how influential instead is the numerical realization. We first show that comparable solution statements do not hold true for Schatten-$p$ functions, including the nuclear norm, and discuss the role of divergent minimizers. Subsequently, we outline corresponding implications for general optimization approaches as well as the more specific IRLS-$0$ algorithm, emphasizing through examples that the transition of the involved smoothing parameter to zero is frequently a more substantial issue than non-convexity. Lastly, we analyze several presented aspects empirically in a series of numerical experiments. In particular, allowing for instance sufficiently many iterations, one may even observe a phase transition for generic recoverability at the absolute theoretical minimum.
</description>
</item>

<item>
<title>
Causal Effect of Functional Treatment
</title>
<link>
http://jmlr.org/papers/v26/23-0381.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0381/23-0381.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Ruoxu Tan, Wei Huang, Zheng Zhang, Guosheng Yin</author>
<description>
We study the causal effect with a functional treatment variable, where practical applications often arise in neuroscience, biomedical sciences, etc. Previous research concerning the effect of a functional variable on an outcome is typically restricted to exploring correlation rather than causality. The generalized propensity score, which is often used to calibrate the selection bias, is not directly applicable to a functional treatment variable due to a lack of definition of probability density function for functional data. We propose three estimators for the average dose-response functional based on the functional linear model, namely, the functional stabilized weight estimator, the outcome regression estimator and the doubly robust estimator, each of which has its own merits. We study their theoretical properties, which are corroborated through extensive numerical experiments. A real data application on electroencephalography data and disease severity demonstrates the practical value of our methods.
</description>
</item>

<item>
<title>
Uplift Model Evaluation with Ordinal Dominance Graphs
</title>
<link>
http://jmlr.org/papers/v26/22-1455.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-1455/22-1455.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Brecht Verbeken, Marie-Anne Guerry, Wouter Verbeke, Sam Verboven</author>
<description>
Uplift modelling is a subfield of causal learning that focuses on ranking entities by individual treatment effects. Uplift models are typically evaluated using Qini curves or Qini scores. While intuitive, the theoretical grounding for Qini in the literature is limited, and the mathematical connection to the well-understood Receiver Operating Characteristic (ROC) curve is unclear. In this paper, we introduce pROCini, a novel uplift evaluation metric that improves upon Qini in two important ways. First, it explicitly incorporates more information by taking into account negative outcomes. Second, it leverages this additional information within the Ordinal Dominance Graph framework, which is the basis behind the well known ROC curve, resulting in a mathematically well-behaved metric that facilitates theoretical analysis. We derive confidence bounds for pROCini, exploiting its theoretical properties. Finally, we empirically validate the improved discriminative power of ROCini and pROCini in a simulation study as well as via experiments on real data.
</description>
</item>

<item>
<title>
High-Dimensional L2-Boosting: Rate of Convergence
</title>
<link>
http://jmlr.org/papers/v26/21-0725.html
</link>
<pdf>
http://jmlr.org/papers/volume26/21-0725/21-0725.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Ye Luo, Martin Spindler, Jannis Kueck</author>
<description>
Boosting is one of the most significant developments in machine learning. This paper studies the rate of convergence of L2-Boosting in a high-dimensional setting under early stopping. We close a gap in the literature and provide the rate of convergence of L2-Boosting in a high-dimensional setting under approximate sparsity and without beta-min condition. We also show that the rate of convergence of the classical L2-Boosting depends on the design matrix described by a sparse eigenvalue condition. To show the latter results, we derive new, improved approximation results for the pure greedy algorithm, based on analyzing the revisiting behavior of L2-Boosting. These results might be of independent interest. Moreover, we introduce so-called  &#34;restricted&#34; L2-Boosting. The restricted L2-Boosting algorithm sticks to the set of the previously chosen variables, exploits the information contained in these variables first and then only occasionally allows to add new variables to this set. We derive the rate of convergence for restricted L2-Boosting under early stopping which is close to the convergence rate of Lasso in an approximate sparse, high-dimensional setting without beta-min condition. We also introduce feasible rules for early stopping, which can be easily implemented and used in applied work. Finally, we present simulation studies to illustrate the relevance of our theoretical results and to provide insights into the practical aspects of boosting. In these simulation studies, L2-Boosting clearly outperforms Lasso. An empirical illustration and the proofs are contained in the Appendix.
</description>
</item>

<item>
<title>
Feature Learning in Finite-Width Bayesian Deep Linear Networks with Multiple Outputs and Convolutional Layers
</title>
<link>
http://jmlr.org/papers/v26/24-1158.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1158/24-1158.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Federico Bassetti, Marco Gherardi, Alessandro Ingrosso, Mauro Pastore, Pietro Rotondo</author>
<description>
Deep linear networks have been extensively studied, as they provide simplified models of deep learning. However, little is known in the case of finite-width architectures with multiple outputs and convolutional layers. In this manuscript, we provide rigorous results for the statistics of functions implemented by the aforementioned class of networks, thus moving closer to a complete characterization of feature learning in the Bayesian setting.  Our results include: (i) an exact and elementary non-asymptotic integral representation for the joint prior distribution over the outputs, given in terms of a mixture of Gaussians; (ii) an analytical formula for the posterior distribution in the case of squared error loss function (Gaussian likelihood); (iii) a quantitative description of the feature learning infinite-width regime, using large deviation theory. From a physical perspective, deep architectures with multiple outputs or convolutional layers represent different manifestations of kernel shape renormalization, and our work provides a dictionary that translates this physics intuition and terminology into rigorous Bayesian statistics.
</description>
</item>

<item>
<title>
How good is your Laplace approximation of the Bayesian posterior? Finite-sample computable error bounds for a variety of useful divergences
</title>
<link>
http://jmlr.org/papers/v26/24-0619.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0619/24-0619.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Mikolaj J. Kasprzak, Ryan Giordano, Tamara Broderick</author>
<description>
The Laplace approximation is a popular method for constructing a Gaussian approximation to the Bayesian posterior and thereby approximating the posterior mean and variance. But approximation quality is a concern. One might consider using rate-of-convergence bounds from certain versions of the Bayesian Central Limit Theorem (BCLT) to provide quality guarantees. But existing bounds require assumptions that are unrealistic even for relatively simple real-life Bayesian analyses; more specifically, existing bounds either (1) require knowing the true data-generating parameter, (2) are asymptotic in the number of samples, (3) do not control the Bayesian posterior mean, or (4) require strongly log concave models to compute. In this work, we provide the first computable bounds on quality that simultaneously (1) do not require knowing the true parameter, (2) apply to finite samples, (3) control posterior means and variances, and (4) apply generally to models that satisfy the conditions of the asymptotic BCLT. Moreover, we substantially improve the dimension dependence of existing bounds; in fact, we achieve the lowest-order dimension dependence possible in the general case. We compute exact constants in our bounds for a variety of standard models, including logistic regression, and numerically demonstrate their utility. We provide a framework for analysis of more complex models.
</description>
</item>

<item>
<title>
Integral Probability Metrics Meet Neural Networks: The Radon-Kolmogorov-Smirnov Test
</title>
<link>
http://jmlr.org/papers/v26/24-0245.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0245/24-0245.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Seunghoon Paik, Michael Celentano, Alden Green, Ryan J. Tibshirani</author>
<description>
Integral probability metrics (IPMs) constitute a general class of nonparametric two-sample tests that are based on maximizing the mean difference between samples from one distribution $P$ versus another $Q$, over all choices of data transformations $f$ living in some function space $\mathcal{F}$. Inspired by recent work that connects what are known as functions of Radon bounded variation (RBV) and neural networks (Parhi and Nowak, 2021, 2023), we study the IPM defined by taking $\mathcal{F}$ to be the unit ball in the RBV space of a given smoothness degree $k \geq 0$. This test, which we refer to as the Radon-Kolmogorov-Smirnov (RKS) test, can be viewed as a generalization of the well-known and classical Kolmogorov-Smirnov (KS) test to multiple dimensions and higher orders of smoothness. It is also intimately connected to neural networks: we prove that the witness in the RKS test—the function $f$ achieving the maximum mean difference—is always a ridge spline of degree $k$, i.e., a single neuron in a neural network. We can thus leverage the power of modern neural network optimization toolkits to (approximately) maximize the criterion that underlies the RKS test. We prove that the RKS test has asymptotically full power at distinguishing any distinct pair $P \not= Q$ of distributions, derive its asymptotic null distribution, and carry out experiments to elucidate the strengths and weaknesses of the RKS test versus the more traditional kernel MMD test.
</description>
</item>

<item>
<title>
On Inference for the Support Vector Machine
</title>
<link>
http://jmlr.org/papers/v26/23-1581.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1581/23-1581.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jakub Rybak, Heather Battey, Wen-Xin Zhou</author>
<description>
The linear support vector machine has a parametrised decision boundary. The paper considers inference for the corresponding parameters, which indicate the effects of individual variables on the decision boundary. The proposed inference is via a convolution-smoothed version of the SVM loss function, this having several inferential advantages over the original SVM, whose associated loss function is not everywhere differentiable. Notably, convolution-smoothing comes with non-asymptotic theoretical guarantees, including a distributional approximation to the parameter estimator that scales more favourably with the dimension of the feature vector. The differentiability of the loss function produces other advantages in some settings; for instance, by facilitating the inclusion of penalties or the synthesis of information from a large number of small samples. The paper closes by relating the linear SVM parameters to those of some probability models for binary outcomes.
</description>
</item>

<item>
<title>
Random Pruning Over-parameterized Neural Networks Can Improve Generalization: A Training Dynamics Analysis
</title>
<link>
http://jmlr.org/papers/v26/23-0832.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0832/23-0832.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Hongru Yang, Yingbin Liang, Xiaojie Guo, Lingfei Wu, Zhangyang Wang</author>
<description>
It has been observed that applying pruning-at-initialization methods and training the sparse networks can sometimes yield slightly better test performance than training the original dense network. Such experimental observations are yet to be understood theoretically. This work makes the first attempt to study this phenomenon. Specifically, we identify a theoretical minimal setting and study a classification task with a one-hidden-layer neural network, which is randomly pruned according to different rates at the initialization. We show that as long as the pruning rate is below a certain threshold, the network provably exhibits good generalization performance after training.More surprisingly, the generalization bound gets better as the pruning rate mildly gets larger. To complement this positive result, we also show a negative result: there exists a large pruning rate such that while gradient descent is still able to drive the training loss toward zero, the generalization performance is no better than random guessing. This further suggests that pruning can change the feature learning process, which leads to the performance drop of the pruned neural network. To our knowledge, this is the first theory work studying how different pruning rates affect neural networks&#39; performance, suggesting that an appropriate pruning rate might improve the neural network&#39;s generalization.
</description>
</item>

<item>
<title>
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
</title>
<link>
http://jmlr.org/papers/v26/23-0058.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0058/23-0058.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang, Aryaman Arora, Zhengxuan Wu, Noah Goodman, Christopher Potts, Thomas Icard</author>
<description>
Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications of the known, but opaque low-level details of black box AI models. Our contributions are (1) generalizing the theory of causal abstraction from mechanism replacement (i.e., hard and soft interventions) to arbitrary mechanism transformation (i.e., functionals from old mechanisms to new mechanisms), (2) providing a flexible, yet precise formalization for the core concepts of polysemantic neurons, the linear representation hypothesis, modular features, and graded faithfulness, and (3) unifying a variety of mechanistic interpretability methods in the common language of causal abstraction, namely, activation and path patching, causal mediation analysis, causal scrubbing, causal tracing, circuit analysis, concept erasure, sparse autoencoders, differential binary masking, distributed alignment search, and steering.
</description>
</item>

<item>
<title>
Implicit vs Unfolded Graph Neural Networks
</title>
<link>
http://jmlr.org/papers/v26/22-0459.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0459/22-0459.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yongyi Yang, Tang Liu, Yangkun Wang, Zengfeng Huang, David Wipf</author>
<description>
It has been observed that message-passing graph neural networks (GNN) sometimes struggle to maintain a healthy balance between the efficient / scalable modeling of long-range dependencies across nodes while avoiding unintended consequences such oversmoothed node representations, sensitivity to spurious edges, or inadequate model interpretability.  To address these and other issues, two separate strategies have recently been proposed, namely implicit and unfolded GNNs (that we abbreviate to IGNN and UGNN respectively).  The former treats node representations as the fixed points of a deep equilibrium model that can efficiently facilitate arbitrary implicit propagation across the graph with a fixed memory footprint.  In contrast, the latter involves treating graph propagation as unfolded descent iterations as applied to some graph-regularized energy function.  While motivated differently, in this paper we carefully quantify explicit situations where the solutions they produce are equivalent and others where their properties sharply diverge.  This includes the analysis of convergence, representational capacity, and interpretability.  In support of this analysis, we also provide empirical head-to-head comparisons across multiple synthetic and public real-world node classification benchmarks.  These results indicate that while IGNN is substantially more memory-efficient, UGNN models support unique, integrated graph attention mechanisms and propagation rules that can achieve strong node classification accuracy across disparate regimes such as adversarially-perturbed graphs, graphs with heterophily, and graphs involving long-range dependencies.
</description>
</item>

<item>
<title>
Towards Optimal Branching of Linear and Semidefinite Relaxations for Neural Network Robustness Certification
</title>
<link>
http://jmlr.org/papers/v26/21-0068.html
</link>
<pdf>
http://jmlr.org/papers/volume26/21-0068/21-0068.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Brendon G. Anderson, Ziye Ma, Jingqi Li, Somayeh Sojoudi</author>
<description>
In this paper, we study certifying the robustness of ReLU neural networks against adversarial input perturbations. To diminish the relaxation error suffered by the popular linear programming (LP) and semidefinite programming (SDP) certification methods, we take a branch-and-bound approach to propose partitioning the input uncertainty set and solving the relaxations on each part separately. We show that this approach reduces relaxation error, and that the error is eliminated entirely upon performing an LP relaxation with a partition intelligently designed to exploit the nature of the ReLU activations. To scale this approach to large networks, we consider using a coarser partition whereby the number of parts in the partition is reduced. We prove that computing such a coarse partition that directly minimizes the LP relaxation error is NP-hard. By instead minimizing the worst-case LP relaxation error, we develop a closed-form branching scheme in the single-hidden layer case. We extend the analysis to the SDP, where the feasible set geometry is exploited to design a branching scheme that minimizes the worst-case SDP relaxation error. Experiments on MNIST, CIFAR-10, and Wisconsin breast cancer diagnosis classifiers demonstrate significant increases in the percentages of test samples certified. By independently increasing the input size and the number of layers, we empirically illustrate under which regimes the branched LP and branched SDP are best applied. Finally, we extend our LP branching method into a multi-layer branching heuristic, which attains comparable performance to prior state-of-the-art heuristics on large-scale, deep neural network certification benchmarks.
</description>
</item>

<item>
<title>
GraphNeuralNetworks.jl: Deep Learning on Graphs with Julia
</title>
<link>
http://jmlr.org/papers/v26/24-2130.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-2130/24-2130.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Carlo Lucibello, Aurora Rossi</author>
<description>
GraphNeuralNetworks.jl is an open-source framework for deep learning on graphs, written in the Julia programming language. It supports multiple GPU backends, generic sparse or dense graph representations, and offers convenient interfaces for manipulating standard, heterogeneous, and temporal graphs with attributes at the node, edge, and graph levels. The framework allows users to define custom graph convolutional layers using gather/scatter message-passing primitives or optimized fused operations. It also includes several popular layers, enabling efficient experimentation with complex deep architectures. The package is available on GitHub: https://github.com/JuliaGraphs/GraphNeuralNetworks.jl.
</description>
</item>

<item>
<title>
Dynamic angular synchronization under smoothness constraints
</title>
<link>
http://jmlr.org/papers/v26/24-0925.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0925/24-0925.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Ernesto Araya, Mihai Cucuringu, Hemant Tyagi</author>
<description>
Given an undirected measurement graph $\mathcal{H} = ([n], \mathcal{E})$, 
the classical angular synchronization problem consists of recovering unknown angles $\theta_1^*,\dots,\theta_n^*$ from a collection of noisy pairwise measurements of the form $(\theta_i^* - \theta_j^*) \mod 2\pi$, for all $\{i,j\} \in \mathcal{E}$. This problem arises in a variety of applications, including computer vision, time synchronization of distributed networks, and ranking from pairwise comparisons. In this paper, we consider a dynamic version of this problem where the angles, and also the measurement graphs evolve over $T$ time points. Assuming a smoothness condition on the evolution of the
latent angles, we derive three algorithms for joint estimation of the angles over all time points. Moreover, for one of the algorithms, we establish non-asymptotic recovery guarantees for the mean-squared error (MSE) under different statistical models. In particular, we show that the MSE converges to zero as $T$ increases under milder conditions than in the static setting. This includes the setting where the measurement graphs are highly sparse and disconnected, and also when the measurement noise is large and can potentially increase with $T$. We complement our theoretical results with experiments on synthetic data.
</description>
</item>

<item>
<title>
Derivative-Informed Neural Operator Acceleration of Geometric MCMC for Infinite-Dimensional Bayesian Inverse Problems
</title>
<link>
http://jmlr.org/papers/v26/24-0745.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0745/24-0745.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Lianghao Cao, Thomas O&#39;Leary-Roseberry, Omar Ghattas</author>
<description>
We propose an operator learning approach to accelerate geometric Markov chain Monte Carlo (MCMC) for solving infinite-dimensional Bayesian inverse problems (BIPs). While geometric MCMC employs high-quality proposals that adapt to posterior local geometry, it requires repeated computations of gradients and Hessians of the log-likelihood, which becomes prohibitive when the parameter-to-observable (PtO) map is defined through expensive-to-solve parametric partial differential equations (PDEs). We consider a delayed-acceptance geometric MCMC method driven by a neural operator surrogate of the PtO map, where the proposal exploits fast surrogate predictions of the log-likelihood and, simultaneously, its gradient and Hessian. To achieve a substantial speedup, the surrogate must accurately approximate the PtO map and its Jacobian, which often demands a prohibitively large number of PtO map samples via conventional operator learning methods. In this work, we present an extension of derivative-informed operator learning [O&#39;Leary-Roseberry et al., J. Comput. Phys., 496 (2024)] that uses joint samples of the PtO map and its Jacobian. This leads to derivative-informed neural operator (DINO) surrogates that accurately predict the observables and posterior local geometry at a significantly lower training cost than conventional methods. Cost and error analysis for reduced basis DINO surrogates are provided. Numerical studies demonstrate that DINO-driven MCMC generates effective posterior samples 3--9 times faster than geometric MCMC and 60--97 times faster than prior geometry-based MCMC. Furthermore, the training cost of DINO surrogates breaks even compared to geometric MCMC after just 10--25 effective posterior samples.
</description>
</item>

<item>
<title>
Wasserstein F-tests for Frechet regression on Bures-Wasserstein manifolds
</title>
<link>
http://jmlr.org/papers/v26/24-0493.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0493/24-0493.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Haoshu Xu, Hongzhe Li</author>
<description>
This paper addresses regression analysis for covariance matrix-valued outcomes with Euclidean covariates, motivated by applications in single-cell genomics and neuroscience where covariance matrices are observed across many samples. Our analysis leverages Fr\&#39;echet regression on the Bures-Wasserstein manifold to estimate the conditional Fr\&#39;echet mean given covariates $x$. We establish a non-asymptotic uniform $\sqrt{n}$-rate of convergence (up to logarithmic factors) over covariates with $\|x\| \lesssim \sqrt{\log n}$ and derive a pointwise central limit theorem to enable statistical inference. For testing covariate effects, we devise a novel test whose null distribution converges to a weighted sum of independent chi-square distributions, with power guarantees against a sequence of contiguous alternatives. Simulations validate the accuracy of the asymptotic theory. Finally, we apply our methods to a single-cell gene expression dataset, revealing age-related changes in gene co-expression networks.
</description>
</item>

<item>
<title>
Distributed Stochastic Bilevel Optimization: Improved Complexity and Heterogeneity Analysis
</title>
<link>
http://jmlr.org/papers/v26/24-0187.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0187/24-0187.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Youcheng Niu, Jinming Xu, Ying Sun, Yan Huang, Li Chai</author>
<description>
This paper considers solving a class of nonconvex-strongly-convex distributed stochastic bilevel optimization (DSBO) problems with personalized inner-level objectives. Most existing algorithms require computational loops for hypergradient estimation, leading to computational inefficiency. Moreover, the impact of data heterogeneity on convergence in bilevel problems is not explicitly characterized yet. To address these issues, we propose LoPA, a loopless personalized distributed algorithm that leverages a tracking mechanism for iterative approximation of inner-level solutions and Hessian-inverse matrices without relying on extra computation loops. Our theoretical analysis explicitly characterizes the heterogeneity across nodes (denoted by $b$), and establishes a sublinear rate of $\mathcal{O}( {\frac{1}{{{{\left( {1 - \rho } \right)}}K}}\!+ \!\frac{{(\frac{b}{\sqrt{m}})^{\frac{2}{3}}  }}{{\left( {1 - \rho } \right)^{\frac{2}{3}} K^{\frac{2}{3}} }} \!+ \!\frac{1}{\sqrt{ K }}( {\sigma _{\operatorname{p} }}  + \frac{1}{\sqrt{m}}{\sigma _{\operatorname{c} }}  ) } )$  without the boundedness of local hypergradients, where ${\sigma _{\operatorname{p} }}$ and ${\sigma _{\operatorname{c} }}$ represent the gradient sampling variances  associated with the inner- and  outer-level variables, respectively.  We also integrate LoPA with a gradient tracking scheme to eliminate the impact of data heterogeneity, yielding an improved rate of ${{\mathcal{O}}}(\frac{{1}}{{ (1-\rho)^2K }} \!+\! \frac{1}{{\sqrt{K}}}( \sigma_{\rm{p}}  \!+\! \frac{1}{\sqrt{m}}\sigma_{\rm{c}} ) )$. The computational complexity of  LoPA is of ${{\mathcal{O}}}({\epsilon^{-2}})$ to an $\epsilon$-stationary point, matching the communication complexity due to the loopless structure, which outperforms existing counterparts for DSBO. Numerical experiments validate the effectiveness of the proposed algorithm. 
</description>
</item>

<item>
<title>
Learning causal graphs via nonlinear sufficient dimension reduction
</title>
<link>
http://jmlr.org/papers/v26/24-0048.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0048/24-0048.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Eftychia Solea, Bing Li, Kyongwon Kim</author>
<description>
We introduce a new nonparametric methodology for estimating a directed acyclic graph (DAG) from observational data. Our method is nonparametric in nature: it does not impose any specific form on the joint distribution of the underlying DAG. Instead, it relies on a linear operator on reproducing kernel Hilbert spaces to evaluate conditional independence. However, a fully nonparametric approach would involve conditioning on a large number of random variables, subjecting it to the curse of dimensionality. To solve this problem, we apply nonlinear sufficient dimension reduction to reduce the number of variables before evaluating the conditional independence. We develop an estimator for the DAG, based on a linear operator that characterizes conditional independence, and establish the consistency and convergence rates of this estimator, as well as the uniform consistency of the estimated Markov equivalence class. We introduce a modified PC-algorithm to implement the estimating procedure efficiently such that the complexity depends on the sparseness of the underlying true DAG. We demonstrate the effectiveness of our methodology through simulations and a real data analysis.
</description>
</item>

<item>
<title>
On Consistent Bayesian Inference from Synthetic Data
</title>
<link>
http://jmlr.org/papers/v26/23-1428.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1428/23-1428.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Ossi Räisä, Joonas Jälkö, Antti Honkela</author>
<description>
Generating synthetic data, with or without differential privacy, has attracted significant attention as a potential solution to the dilemma between making data easily available, and the privacy of data subjects. Several works have shown that consistency of downstream analyses from synthetic data, including accurate uncertainty estimation, requires accounting for the synthetic data generation. There are very few methods of doing so, most of them for frequentist analysis. In this paper, we study how to perform consistent Bayesian inference from synthetic data. We prove that mixing posterior samples obtained separately from multiple large synthetic data sets, that are sampled from a posterior predictive, converges to the posterior of the downstream analysis under standard regularity conditions when the analyst&#39;s model is compatible with the data provider&#39;s model. We also present several examples showing how the theory works in practice, and showing how Bayesian inference can fail when the compatibility assumption is not met, or the synthetic data set is not significantly larger than the original.
</description>
</item>

<item>
<title>
Optimization Over a Probability Simplex
</title>
<link>
http://jmlr.org/papers/v26/23-1166.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1166/23-1166.pdf
</pdf>
<pubDate>2025</pubDate>
<author>James Chok, Geoffrey M. Vasil</author>
<description>
We propose a new iteration scheme, the Cauchy-Simplex, to optimize convex problems over the probability simplex $\{w\in\mathbb{R}^n\ |\ \sum_i w_i=1\ \textrm{and}\ w_i\geq0\}$.
Specifically, we map the simplex to the positive quadrant of a unit sphere, envisage gradient descent in latent variables, and map the result back in a way that only depends on the simplex variable. Moreover, proving rigorous convergence results in this formulation leads inherently to tools from information theory (e.g., cross-entropy and KL divergence). Each iteration of the Cauchy-Simplex consists of simple operations, making it well-suited for high-dimensional problems. In continuous time, we prove that $f(x_T)-f(x^*) = O(1/T)$ for differentiable real-valued convex functions, where $T$ is the number of time steps and $w^*$ is the optimal solution. Numerical experiments of projection onto convex hulls show faster convergence than similar algorithms. Finally, we apply our algorithm to online learning problems and prove the convergence of the average regret for (1) Prediction with expert advice and (2) Universal Portfolios.
</description>
</item>

<item>
<title>
Laplace Meets Moreau: Smooth Approximation to Infimal Convolutions Using Laplace&#39;s Method
</title>
<link>
http://jmlr.org/papers/v26/24-0944.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0944/24-0944.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Ryan J. Tibshirani, Samy Wu Fung, Howard Heaton, Stanley Osher</author>
<description>
We study approximations to the Moreau envelope---and infimal convolutions more broadly---based on Laplace&#39;s method, a classical tool in analysis which ties certain integrals to suprema of their integrands. We believe the connection between Laplace&#39;s method and infimal convolutions is generally deserving of more attention in the study of optimization and partial differential equations, since it bears numerous potentially important applications, from proximal-type algorithms to Hamilton-Jacobi equations.
</description>
</item>

<item>
<title>
Sampling and Estimation on Manifolds using the Langevin Diffusion
</title>
<link>
http://jmlr.org/papers/v26/24-0829.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0829/24-0829.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Karthik Bharath, Alexander Lewis, Akash Sharma, Michael V. Tretyakov</author>
<description>
Error bounds are derived for sampling and estimation using a discretization of an intrinsically defined Langevin diffusion with invariant measure $\text{d}\mu_\phi \propto e^{-\phi} \mathrm{dvol}_g $ on a compact Riemannian manifold.  Two estimators of linear functionals of $\mu_\phi $ based on the discretized Markov process are considered: a time-averaging estimator based on a single trajectory and an ensemble-averaging estimator based on multiple independent trajectories. Imposing no restrictions beyond a nominal level of smoothness on $\phi$, first-order error bounds, in discretization step size, on the bias and variance/mean-square error of both estimators are derived. The order of error matches the optimal rate in Euclidean and flat spaces, and leads to a first-order bound on distance between the invariant measure $\mu_\phi$ and a stationary measure of the discretized Markov process. This order is preserved even upon using retractions when exponential maps are unavailable in closed form, thus enhancing practicality of the proposed algorithms. Generality of the proof techniques, which exploit links between two partial differential equations and the semigroup of operators corresponding to the Langevin diffusion, renders them amenable for the study of a more general class of sampling algorithms related to the Langevin diffusion. Conditions for extending analysis to the case of non-compact manifolds are discussed. Numerical illustrations with distributions, log-concave and otherwise, on the manifolds of positive and negative curvature elucidate on the derived bounds and demonstrate practical utility of the sampling algorithm.
</description>
</item>

<item>
<title>
Sharp Bounds for Sequential Federated Learning on Heterogeneous Data
</title>
<link>
http://jmlr.org/papers/v26/24-0668.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0668/24-0668.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yipeng Li, Xinchen Lyu</author>
<description>
There are two paradigms in Federated Learning (FL): parallel FL (PFL), where models are trained in a parallel manner across clients, and sequential FL (SFL), where models are trained in a sequential manner across clients. Specifically, in PFL, clients perform local updates independently and send the updated model parameters to a global server for aggregation; in SFL, one client starts its local updates only after receiving the model parameters from the previous client in the sequence. In contrast to that of PFL, the convergence theory of SFL on heterogeneous data is still lacking. To resolve the theoretical dilemma of SFL, we establish sharp convergence guarantees for SFL on heterogeneous data with both upper and lower bounds. Specifically, we derive the upper bounds for the strongly convex, general convex and non-convex objective functions, and construct the matching lower bounds for the strongly convex and general convex objective functions. Then, we compare the upper bounds of SFL with those of PFL, showing that SFL outperforms PFL on heterogeneous data (at least, when the level of heterogeneity is relatively high). Experimental results validate the counterintuitive theoretical finding.
</description>
</item>

<item>
<title>
Local Linear Recovery Guarantee of Deep Neural Networks at Overparameterization
</title>
<link>
http://jmlr.org/papers/v26/24-0192.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0192/24-0192.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yaoyu Zhang, Leyang Zhang, Zhongwang Zhang, Zhiwei Bai</author>
<description>
Determining whether deep neural network (DNN) models can reliably recover target functions at overparameterization is a critical yet complex issue in the theory of deep learning. To advance understanding in this area, we introduce a concept we term “local linear recovery” (LLR), a weaker form of target function recovery that renders the problem more amenable to theoretical analysis. In the sense of LLR, we prove that functions expressible by narrower DNNs are guaranteed to be recoverable from fewer samples than model parameters. Specifically, we establish upper limits on the optimistic sample sizes, defined as the smallest sample size necessary to guarantee LLR, for functions in the space of a given DNN. Furthermore, we prove that these upper bounds are achieved in the case of two-layer tanh neural networks. Our research lays a solid groundwork for future investigations into the recovery capabilities of DNNs in overparameterized scenarios.
</description>
</item>

<item>
<title>
Stabilizing Sharpness-Aware Minimization Through A Simple Renormalization Strategy
</title>
<link>
http://jmlr.org/papers/v26/24-0065.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0065/24-0065.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Chengli Tan, Jiangshe Zhang, Junmin Liu, Yicheng Wang, Yunda Hao</author>
<description>
Recently, sharpness-aware minimization (SAM) has attracted much attention because of its surprising effectiveness in improving generalization performance. However, compared to stochastic gradient descent (SGD), it is more prone to getting stuck at the saddle points, which as a result may lead to performance degradation. To address this issue, we propose a simple renormalization strategy, dubbed Stable SAM (SSAM), so that the gradient norm of the descent step maintains the same as that of the ascent step. Our strategy is easy to implement and flexible enough to integrate with SAM and its variants, almost at no computational cost. With elementary tools from convex optimization and learning theory, we also conduct a theoretical analysis of sharpness-aware training, revealing that compared to SGD, the effectiveness of SAM is only assured in a limited regime of learning rate. In contrast, we show how SSAM extends this regime of learning rate and then it can consistently perform better than SAM with the minor modification. Finally, we demonstrate the improved performance of SSAM on several representative data sets and tasks.
</description>
</item>

<item>
<title>
Fine-Grained Change Point Detection for Topic Modeling with Pitman-Yor Process
</title>
<link>
http://jmlr.org/papers/v26/23-1576.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1576/23-1576.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Feifei Wang, Zimeng Zhao, Ruimin Ye, Xiaoge Gu, Xiaoling Lu</author>
<description>
Identifying change points in dynamic text data is crucial for understanding the evolving nature of topics across various sources, such as news articles, scientific papers, and social media posts. While topic modeling has become a widely used technique for this purpose, capturing fine-grained shifts in individual topics over time remains a significant challenge. Traditional approaches typically use a two-stage process, separating topic modeling and change point detection. However, this separation can lead to information loss and inconsistency in capturing subtle changes in topic evolution. To address this issue, we propose TOPIC-PYP, a change point detection model specifically designed for fine-grained topic-level analysis, i.e., detecting change points for each individual topic. By leveraging the Pitman-Yor process, TOPIC-PYP effectively captures the dynamic evolution of topic meanings over time. Unlike traditional methods, TOPIC-PYP integrates topic modeling and change point detection into a unified framework, facilitating a more comprehensive understanding of the relationship between topic evolution and change points. Experimental evaluations on both synthetic and real-world datasets demonstrate the effectiveness of TOPIC-PYP in accurately detecting change points and generating high-quality topics.
</description>
</item>

<item>
<title>
Deletion Robust Non-Monotone Submodular Maximization over Matroids
</title>
<link>
http://jmlr.org/papers/v26/23-1219.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1219/23-1219.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Paul Dütting, Federico Fusco, Silvio Lattanzi, Ashkan Norouzi-Fard, Morteza Zadimoghaddam</author>
<description>
We study the deletion robust version of submodular maximization under matroid constraints. The goal is to extract a small-size summary of the data set that contains a high-value independent set even after an adversary deletes some elements. We present constant-factor approximation algorithms, whose space complexity depends on the rank $k$ of the matroid, the number $d$ of deleted elements, and the input precision $\varepsilon$. In the centralized setting we present a $(4.494+O(\varepsilon))$-approximation algorithm with summary size $O( \frac{k+d}{\varepsilon^2}\log \frac{k}{\varepsilon})$ that improves to a $(3.582+O(\varepsilon))$-approximation with $O(k + \frac{d}{\varepsilon^2}\log \frac{k}{\varepsilon})$ summary size when the objective is monotone.  In the streaming setting we provide a $(9.294 + O(\varepsilon))$-approximation algorithm with summary size and memory $O(k + \frac{d}{\varepsilon^2}\log \frac{k}{\varepsilon})$; the approximation factor is then improved to  $(5.582+O(\varepsilon))$ in the monotone case.
</description>
</item>

<item>
<title>
Instability, Computational Efficiency and Statistical Accuracy
</title>
<link>
http://jmlr.org/papers/v26/22-0300.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0300/22-0300.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Nhat Ho, Koulik Khamaru, Raaz Dwivedi, Martin J. Wainwright, Michael I. Jordan, Bin Yu</author>
<description>
Many statistical estimators are defined as the fixed point of a data-dependent operator, with estimators based on minimizing a cost function being an important special case.  The limiting performance of such estimators depends on the properties of the population-level operator in the idealized limit of infinitely many samples.  We develop a general framework that yields bounds on statistical accuracy based on the interplay between the deterministic convergence rate of the algorithm at the population level, and its degree of (in)stability when applied to an empirical object based on $n$ samples.  Using this framework, we analyze both stable forms of gradient descent and some higher-order and unstable algorithms, including Newton&#39;s method and its cubic-regularized variant, as well as the EM algorithm. We provide applications of our general results to several concrete classes of models, including Gaussian mixture estimation, non-linear regression models, and informative non-response models.  We exhibit cases in which an unstable algorithm can achieve the same statistical accuracy as a stable algorithm in exponentially fewer steps---namely, with the number of iterations being reduced from polynomial to logarithmic in sample size $n$.
</description>
</item>

<item>
<title>
Estimation of Local Geometric Structure on Manifolds from Noisy Data
</title>
<link>
http://jmlr.org/papers/v26/25-0183.html
</link>
<pdf>
http://jmlr.org/papers/volume26/25-0183/25-0183.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yariv Aizenbud, Barak Sober</author>
<description>
A common observation in data-driven applications is that high-dimensional data have a low intrinsic dimension, at least locally. In this work, we consider the problem of point estimation for manifold-valued data. Namely, given a finite set of noisy samples of $\mathcal{M}$, a $d$ dimensional submanifold of $\mathbb{R}^D$, and a point $r$ near the manifold we aim to project $r$ onto the manifold. Assuming that the data was sampled uniformly from a tubular neighborhood of a $k$-times smooth boundaryless and compact manifold, we present an algorithm that takes $r$ from this neighborhood and outputs $\hat p_n\in \mathbb{R}^D$, and $\widehat{T_{\hat p_n}\mathcal{M}}$ an element in the Grassmannian $Gr(d, D)$. We prove that as the number of samples $n\to\infty$, the point $\hat p_n$ converges to $\mathbf{p}\in \mathcal{M}$, the projection of $r$ onto $\mathcal{M}$, and $\widehat{T_{\hat p_n}\mathcal{M}}$ converges to $T_{\mathbf{p}}\mathcal{M}$ (the tangent space at that point) with high probability. Furthermore, we show that $\hat p_n$ approaches the manifold with an asymptotic rate of $n^{-\frac{k}{2k + d}}$, and that $\hat p_n, \widehat{T_{\hat p_n}\mathcal{M}}$ approach $\mathbf{p}$ and $T_{\mathbf{p}}\mathcal{M}$ correspondingly with asymptotic rates of $n^{-\frac{k-1}{2k + d}}$. %While we These rates coincide with the optimal rates for the estimation of function derivatives.
</description>
</item>

<item>
<title>
Ontolearn---A Framework for Large-scale OWL Class Expression Learning in Python
</title>
<link>
http://jmlr.org/papers/v26/24-1113.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1113/24-1113.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Caglar Demir, Alkid Baci, N&#39;Dah Jean Kouagou, Leonie Nora Sieger, Stefan Heindorf, Simon Bin, Lukas Blübaum, Alexander Bigerl, Axel-Cyrille Ngonga Ngomo</author>
<description>
In this paper, we present Ontolearn---a framework for learning OWL class expressions over large knowledge graphs.
Ontolearn contains efficient implementations of recent state-of-the-art symbolic and neuro-symbolic class expression learners including EvoLearner and DRILL.
A learned OWL class expression can be used to classify instances in the knowledge graph.
Furthermore, Ontolearn integrates a verbalization module based on an LLM to translate complex OWL class expressions into natural language sentences.
By mapping OWL class expressions into respective SPARQL queries, Ontolearn can be easily used to operate over a remote triplestore.
The source code of Ontolearn is available at https://github.com/dice-group/Ontolearn.
</description>
</item>

<item>
<title>
Continuously evolving rewards in an open-ended environment
</title>
<link>
http://jmlr.org/papers/v26/24-0847.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0847/24-0847.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Richard M. Bailey</author>
<description>
Unambiguous identification of the rewards driving behaviours of entities operating in complex open-ended real-world environments is difficult, in part because goals and associated behaviours emerge endogenously and are dynamically updated as environments change. Reproducing such dynamics in models would be useful in many domains, particularly where fixed reward functions limit the adaptive capabilities of agents. Simulation experiments described here assess a candidate algorithm for the dynamic updating of the reward function, RULE: Reward Updating through Learning and Expectation. The approach is tested in a simplified ecosystem-like setting where experiments challenge entities&#39; survival, calling for significant behavioural change. The population of entities successfully demonstrate the abandonment of an initially rewarded but ultimately detrimental behaviour, amplification of beneficial behaviour, and appropriate responses to novel items added to their environment. These adjustments happen through endogenous modification of the entities&#39; reward function, during continuous learning, without external intervention.
</description>
</item>

<item>
<title>
Recursive Causal Discovery
</title>
<link>
http://jmlr.org/papers/v26/24-0384.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0384/24-0384.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Ehsan Mokhtarian, Sepehr Elahi, Sina Akbari, Negar Kiyavash</author>
<description>
Causal discovery from observational data, i.e., learning the causal graph from a finite set of samples from the joint distribution of the variables, is often the first step toward the identification and estimation of causal effects, a key requirement in numerous scientific domains. Causal discovery is hampered by two main challenges: limited data results in errors in statistical testing and the computational complexity of the learning task is daunting. This paper builds upon and extends four of our prior publications (Mokhtarian et al., 2021; Akbari et al., 2021; Mokhtarian et al., 2022, 2023a). These works introduced the concept of removable variables, which are the only variables that can be removed recursively for the purpose of causal discovery. Presence and identification of removable variables allow recursive approaches for causal discovery, a promising solution that helps to address the aforementioned challenges by reducing the problem size successively. This reduction not only minimizes conditioning sets in each conditional independence (CI) test, leading to fewer errors but also significantly decreases the number of required CI tests. The worst-case performances of these methods nearly match the lower bound. In this paper, we present a unified framework for the proposed algorithms, refined with additional details and enhancements for a coherent presentation. A comprehensive literature review is also included, comparing the computational complexity of our methods with existing approaches, showcasing their state-of-the-art efficiency. Another contribution of this paper is the release of RCD, a Python package that efficiently implements these algorithms. This package is designed for practitioners and researchers interested in applying these methods in practical scenarios. The package is available at github.com/ban-epfl/rcd, with comprehensive documentation provided at rcdpackage.com.
</description>
</item>

<item>
<title>
Evaluation of Active Feature Acquisition Methods for Time-varying Feature Settings
</title>
<link>
http://jmlr.org/papers/v26/23-1635.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1635/23-1635.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Henrik von Kleist, Alireza Zamanian, Ilya Shpitser, Narges Ahmidi</author>
<description>
Machine learning methods often assume that input features are available at no cost. However, in domains like healthcare, where acquiring features could be expensive or harmful, it is necessary to balance a feature&#39;s acquisition cost against its predictive value. The task of training an AI agent to decide which features to acquire is called active feature acquisition (AFA). By deploying an AFA agent, we effectively alter the acquisition strategy and trigger a distribution shift. To safely deploy AFA agents under this distribution shift, we present the problem of active feature acquisition performance evaluation (AFAPE). We examine AFAPE under i) a no direct effect (NDE) assumption, stating that acquisitions do not affect the underlying feature values; and ii) a no unobserved confounding (NUC) assumption, stating that retrospective feature acquisition decisions were only based on observed features. We show that one can apply missing data methods under the NDE assumption and offline reinforcement learning under the NUC assumption. When NUC and NDE hold, we propose a novel semi-offline reinforcement learning framework. This framework requires a weaker positivity assumption and introduces three new estimators: A direct method (DM), an inverse probability weighting (IPW), and a double reinforcement learning (DRL) estimator.
</description>
</item>

<item>
<title>
On Adaptive Stochastic Optimization for Streaming Data: A Newton&#39;s Method with O(dN) Operations
</title>
<link>
http://jmlr.org/papers/v26/23-1565.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1565/23-1565.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Antoine Godichon-Baggioni, Nicklas Werge</author>
<description>
Stochastic optimization methods face new challenges in the realm of streaming data, characterized by a continuous flow of large, high-dimensional data. While first-order methods, like stochastic gradient descent, are the natural choice for such data, they often struggle with ill-conditioned problems. In contrast, second-order methods, such as Newton&#39;s method, offer a potential solution but are computationally impractical for large-scale streaming applications. This paper introduces adaptive stochastic optimization methods that effectively address ill-conditioned problems while functioning in a streaming context. Specifically, we present adaptive inversion-free stochastic quasi-Newton methods with computational complexity matching that of first-order methods, $\mathcal{O}(dN)$, where $d$ represents the number of dimensions/features and $N$ the number of data points. Theoretical analysis establishes their asymptotic efficiency, and empirical studies demonstrate their effectiveness in scenarios with complex covariance structures and poor initializations. In particular, we demonstrate that our adaptive quasi-Newton methods can outperform or match existing first- and second-order methods.
</description>
</item>

<item>
<title>
Determine the Number of States in Hidden Markov Models via Marginal Likelihood
</title>
<link>
http://jmlr.org/papers/v26/23-0343.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0343/23-0343.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yang Chen, Cheng-Der Fuh, Chu-Lan Michael Kao</author>
<description>
Hidden Markov models (HMM) have been widely used by scientists to model stochastic systems: the underlying process is a discrete Markov chain, and the observations are noisy realizations of the underlying process. Determining the number of hidden states for an HMM is a model selection problem which is yet to be satisfactorily solved, especially for the popular Gaussian HMM with heterogeneous covariance. In this paper, we propose a consistent method for determining the number of hidden states of HMM based on the marginal likelihood, which is obtained by integrating out both the parameters and hidden states. Moreover, we show that the model selection problem of HMM includes the order selection problem of finite mixture models as a special case. We give rigorous proof of the consistency of the proposed marginal likelihood method and provide an efficient computation method for practical implementation. We numerically compare the proposed method with the Bayesian information criterion (BIC), demonstrating the effectiveness of the proposed marginal likelihood method.
</description>
</item>

<item>
<title>
Variance-Aware Estimation of Kernel Mean Embedding
</title>
<link>
http://jmlr.org/papers/v26/23-0161.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0161/23-0161.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Geoffrey Wolfer, Pierre Alquier</author>
<description>
An important feature of kernel mean embeddings (KME) is that the rate of convergence of the empirical KME to the true distribution KME can be bounded independently of the dimension of the space, properties of the distribution and smoothness features of the kernel. We show how to speed-up convergence by leveraging variance information in the reproducing kernel Hilbert space. Furthermore, we show that even when such information is a priori unknown, we can efficiently estimate it from the data, recovering the desiderata of a distribution agnostic bound that enjoys acceleration in fortuitous settings. We further extend our results from independent data to stationary mixing sequences and illustrate our methods in the context of hypothesis testing and robust parametric estimation.
</description>
</item>

<item>
<title>
Scaling ResNets in the Large-depth Regime
</title>
<link>
http://jmlr.org/papers/v26/22-0664.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0664/22-0664.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Pierre Marion, Adeline Fermanian, Gérard Biau, Jean-Philippe Vert</author>
<description>
Deep ResNets are recognized for achieving state-of-the-art results in complex machine learning tasks. However, the remarkable performance of these architectures relies on a training procedure that needs to be carefully crafted to avoid vanishing or exploding gradients, particularly as the depth $L$ increases. No consensus has been reached on how to mitigate this issue, although a widely discussed strategy consists in scaling the output of each layer by a factor $\alpha_L$. We show in a probabilistic setting that with standard i.i.d. initializations, the only non-trivial dynamics is for $\alpha_L = \frac{1}{\sqrt{L}}$---other choices lead either to explosion or to identity mapping. This scaling factor corresponds in the continuous-time limit to a neural stochastic differential equation, contrarily to a widespread interpretation that deep ResNets are discretizations of neural ordinary differential equations. By contrast, in the latter regime, stability is obtained with specific correlated initializations and $\alpha_L = \frac{1}{L}$. Our analysis suggests a strong interplay between scaling and regularity of the weights as a function of the layer index. Finally, in a series of experiments, we exhibit a continuous range of regimes driven by these two parameters, which jointly impact performance before and after training.
</description>
</item>

<item>
<title>
A Comparative Evaluation of Quantification Methods
</title>
<link>
http://jmlr.org/papers/v26/21-0241.html
</link>
<pdf>
http://jmlr.org/papers/volume26/21-0241/21-0241.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Tobias Schumacher, Markus Strohmaier, Florian Lemmerich</author>
<description>
Quantification represents the problem of estimating the distribution of class labels on unseen data. It also represents a growing research field in supervised machine learning, for which a large variety of different algorithms has been proposed in recent years. However, a comprehensive empirical comparison of quantification methods that supports algorithm selection is not available yet. In this work, we close this research gap by conducting a thorough empirical performance comparison of 24 different quantification methods on in total more than 40 datasets, considering binary as well as multiclass quantification settings. We observe that no single algorithm generally outperforms all competitors, but identify a group of methods that perform best in the binary setting, including the threshold selection-based median sweep and TSMax methods, the DyS framework including the HDy method, Forman&#39;s mixture model, and Friedman&#39;s method. For the multiclass setting, we observe that a different, broad group of algorithms yields good performance, including the HDx method, the generalized probabilistic adjusted count, the readme method, the energy distance minimization method, the EM algorithm for quantification, and Friedman&#39;s method. We also find that tuning the underlying classifiers has in most cases only a limited impact on the quantification performance. More generally, we find that the performance on multiclass quantification is inferior to the results obtained in the binary setting. Our results can guide practitioners who intend to apply quantification algorithms and help researchers identify opportunities for future research.
</description>
</item>

<item>
<title>
Lightning UQ Box: Uncertainty Quantification for Neural Networks
</title>
<link>
http://jmlr.org/papers/v26/24-2110.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-2110/24-2110.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Nils Lehmann, Nina Maria Gottschling, Jakob Gawlikowski, Adam J. Stewart, Stefan Depeweg, Eric Nalisnick</author>
<description>
Although neural networks have shown impressive results in a multitude of application domains, the &#34;black box&#34; nature of deep learning and lack of confidence estimates have led to scepticism, especially in domains like medicine and physics where such estimates are critical. Research on uncertainty quantification (UQ) has helped elucidate the reliability of these models, but existing implementations of these UQ methods are sparse and difficult to reuse. To this end, we introduce Lightning UQ Box, a PyTorch-based Python library for deep learning-based UQ methods powered by PyTorch Lightning. Lightning UQ Box supports classification, regression, semantic segmentation, and pixelwise regression applications, and UQ methods from a variety of theoretical motivations. With this library, we provide an entry point for practitioners new to UQ, as well as easy-to-use components and tools for scalable deep learning applications.
</description>
</item>

<item>
<title>
Scaling Data-Constrained Language Models
</title>
<link>
http://jmlr.org/papers/v26/24-1000.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1000/24-1000.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Niklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, Colin Raffel</author>
<description>
The current trend of scaling language models involves increasing both parameter count and training data set size. Extrapolating this trend suggests that training data set size may soon be limited by the amount of text data available on the internet. Motivated by this limit, we investigate scaling language models in data-constrained regimes. Specifically, we run a large set of experiments varying the extent of data repetition and compute budget, ranging up to 900 billion training tokens and 9 billion parameter models. We find that with constrained data for a fixed compute budget, training with up to 4 epochs of repeated data yields negligible changes to loss compared to having unique data. However, with more repetition, the value of adding compute eventually decays to zero. We propose and empirically validate a scaling law for compute optimality that accounts for the decreasing value of repeated tokens and excess parameters. Finally, we experiment with approaches mitigating data scarcity, including augmenting the training data set with code data or removing commonly used filters. Models and data sets from our 400 training runs are freely available at https://github.com/huggingface/datablations.
</description>
</item>

<item>
<title>
Curvature-based Clustering on Graphs
</title>
<link>
http://jmlr.org/papers/v26/24-0781.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0781/24-0781.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yu Tian, Zachary Lubberts, Melanie Weber</author>
<description>
Unsupervised node clustering (or community detection) is a classical graph learning task. In this paper, we study algorithms that exploit the geometry of the graph to identify densely connected substructures, which form clusters or communities. Our method implements discrete Ricci curvatures and their associated geometric flows, under which the edge weights of the graph evolve to reveal its community structure. We consider several discrete curvature notions and analyze the utility of the resulting algorithms. In contrast to prior literature, we study not only single-membership community detection, where each node belongs to exactly one community, but also mixed-membership community detection, where communities may overlap. For the latter, we argue that it is beneficial to perform community detection on the line graph, i.e., the graph&#39;s dual. We provide both theoretical and empirical evidence for the utility of our curvature-based clustering algorithms. In addition, we give several results on the relationship between the curvature of a graph and that of its dual, which enable the efficient implementation of our proposed mixed-membership community detection approach and which may be of independent interest for curvature-based network analysis.
</description>
</item>

<item>
<title>
Composite Goodness-of-fit Tests with Kernels
</title>
<link>
http://jmlr.org/papers/v26/24-0276.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0276/24-0276.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Oscar Key, Arthur Gretton, François-Xavier Briol, Tamara Fernandez</author>
<description>
We propose kernel-based hypothesis tests for the challenging composite testing problem, where we are interested in whether the data comes from any distribution in some parametric family. Our tests make use of minimum distance estimators based on kernel-based distances such as the maximum mean discrepancy. As our main result, we show that we are able to estimate the parameter and conduct our test on the same data (without data splitting), while maintaining a correct test level. We also prove that the popular wild bootstrap will lead to an overly conservative test, and show that the parametric bootstrap is consistent and can lead to significantly improved performance in practice. Our approach is illustrated on a range of problems, including testing for goodness-of-fit of a non-parametric density model, and an intractable generative model of a biological cellular network.
</description>
</item>

<item>
<title>
PFLlib: A Beginner-Friendly and Comprehensive Personalized Federated Learning Library and Benchmark
</title>
<link>
http://jmlr.org/papers/v26/23-1634.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1634/23-1634.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jianqing Zhang, Yang Liu, Yang Hua, Hao Wang, Tao Song, Zhengui Xue, Ruhui Ma, Jian Cao</author>
<description>
Amid the ongoing advancements in Federated Learning (FL), a machine learning paradigm that allows collaborative learning with data privacy protection, personalized FL (pFL) has gained significant prominence as a research direction within the FL domain. Whereas traditional FL (tFL) focuses on jointly learning a global model, pFL aims to balance each client&#39;s global and personalized goals in FL settings. To foster the pFL research community, we started and built PFLlib, a comprehensive pFL library with an integrated benchmark platform. In PFLlib, we implemented 37 state-of-the-art FL algorithms (8 tFL algorithms and 29 pFL algorithms) and provided various evaluation environments with three statistically heterogeneous scenarios and 24 datasets. At present, PFLlib has gained more than 1600 stars and 300 forks on GitHub.
</description>
</item>

<item>
<title>
The Effect of SGD Batch Size on Autoencoder Learning: Sparsity, Sharpness, and Feature Learning
</title>
<link>
http://jmlr.org/papers/v26/23-1022.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1022/23-1022.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Nikhil Ghosh, Spencer Frei, Wooseok Ha, Bin Yu</author>
<description>
In this work, we investigate the dynamics of stochastic gradient descent (SGD) when training a single-neuron autoencoder with linear or ReLU activation on orthogonal data. We show that for this non-convex problem, randomly initialized SGD with a constant step size successfully finds a global minimum for any batch size choice. However, the particular global minimum found depends upon the batch size. In the full-batch setting, we show that the solution is dense (i.e., not sparse) and is highly aligned with its initialized direction, showing that relatively little feature learning occurs. On the other hand, for any batch size strictly smaller than the number of samples, SGD finds a global minimum that is sparse and nearly orthogonal to its initialization, showing that the randomness of stochastic gradients induces a qualitatively different type of &#34;feature selection&#34; in this setting. Moreover, if we measure the sharpness of the minimum by the trace of the Hessian, the minima found with full-batch gradient descent are flatter than those found with strictly smaller batch sizes, in contrast to previous works which suggest that large batches lead to sharper minima. To prove convergence of SGD with a constant step size, we introduce a powerful tool from the theory of non-homogeneous random walks which may be of independent interest.
</description>
</item>

<item>
<title>
Efficient and Robust Transfer Learning of Optimal Individualized Treatment Regimes with Right-Censored Survival Data
</title>
<link>
http://jmlr.org/papers/v26/23-0335.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0335/23-0335.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Pan Zhao, Julie Josse, Shu Yang</author>
<description>
An individualized treatment regime (ITR) is a decision rule that assigns treatments based on patients&#39; characteristics. The value function of an ITR is the expected outcome in a counterfactual world had this ITR been implemented. Recently, there has been increasing interest in combining heterogeneous data sources, such as leveraging the complementary features of randomized controlled trial (RCT) data and a large observational study (OS). Usually, a covariate shift exists between the source and target population, rendering the source-optimal ITR not optimal for the target population. We present an efficient and robust transfer learning framework for estimating the optimal ITR with right-censored survival data that generalizes well to the target population. The value function accommodates a broad class of functionals of survival distributions, including survival probabilities and restrictive mean survival times (RMSTs). We propose a doubly robust estimator of the value function, and the optimal ITR is learned by maximizing the value function within a pre-specified class of ITRs. We establish the cubic rate of convergence for the estimated parameter indexing the optimal ITR, and show that the proposed optimal value estimator is consistent and asymptotically normal even with flexible machine learning methods for nuisance parameter estimation. We evaluate the empirical performance of the proposed method by simulation studies and a real data application of sodium bicarbonate therapy for patients with severe metabolic acidaemia in the intensive care unit (ICU), combining a RCT and an observational study with heterogeneity.
</description>
</item>

<item>
<title>
DAGs as Minimal I-maps for the Induced Models of Causal Bayesian Networks under Conditioning
</title>
<link>
http://jmlr.org/papers/v26/23-0002.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0002/23-0002.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Xiangdong Xie, Jiahua Guo, Yi Sun</author>
<description>
Bayesian networks (BNs) are a powerful tool for knowledge representation and reasoning, especially for complex systems.  A critical task in the applications of BNs is conditional inference or inference in the presence of selection bias. However, post-conditioning, the conditional distribution family of a BN can become complex for analysis, and the corresponding induced subgraph may not accurately encode the  conditional independencies  for the remaining variables. In this work, we first investigate the conditions under which a BN remains closed under conditioning, meaning that the induced subgraph is consistent with the structural information of conditional distributions. Conversely, when a BN is not closed, we aim to construct a new directed acyclic graph (DAG) as a minimal $\mathcal{I}$-map for the conditional model by incorporating directed edges into the original induced graph. We present an equivalent characterization of this minimal $\mathcal{I}$-map and develop an efficient algorithm for its identification.  The proposed framework improves the efficiency of conditional inference of a BN.  Additionally, the DAG minimal $\mathcal{I}$-map offers graphical criteria for the safe integration of knowledge from diverse sources (subpopulations/conditional distributions), facilitating correct parameter estimation. Both theoretical analysis and simulation studies demonstrate that using a DAG minimal $\mathcal{I}$-map for conditional inference is more effective than traditional methods based on the joint distribution of the original BN.
</description>
</item>

<item>
<title>
Adjusted Expected Improvement for Cumulative Regret Minimization in Noisy Bayesian Optimization
</title>
<link>
http://jmlr.org/papers/v26/22-0523.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0523/22-0523.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Shouri Hu, Haowei Wang, Zhongxiang Dai, Bryan Kian Hsiang Low, Szu Hui Ng</author>
<description>
The expected improvement (EI) is one of the most popular acquisition functions for Bayesian optimization (BO) and has demonstrated good empirical performances in many applications for the minimization of simple regret. However, under the evaluation metric of cumulative regret, the performance of EI may not be competitive, and its existing theoretical regret upper bound still has room for improvement. To adapt the EI for better performance under cumulative regret, we introduce a novel quantity called the evaluation cost which is compared against the acquisition function, and with this, develop the expected improvement-cost (EIC) algorithm. In each iteration of EIC, a new point with the largest acquisition function value is sampled, only if that value exceeds its evaluation cost. If none meets this criteria, the current best point is resampled. This evaluation cost quantifies the potential downside of sampling a point, which is important under the cumulative regret metric as the objective function value in every iteration affects the performance measure. We establish in theory a high-probability regret upper bound of EIC based on the maximum information gain, which is tighter than the bound of existing EI-based algorithms. It is also comparable to the regret bound of other popular BO algorithms such as Thompson sampling (GP-TS) and upper confidence bound (GP-UCB). We further perform experiments to illustrate the improvement of EIC over several popular BO algorithms.
</description>
</item>

<item>
<title>
Manifold Fitting under Unbounded Noise
</title>
<link>
http://jmlr.org/papers/v26/21-0039.html
</link>
<pdf>
http://jmlr.org/papers/volume26/21-0039/21-0039.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Zhigang Yao, Yuqing Xia</author>
<description>
In the field of non-Euclidean statistical analysis, a trend has emerged in recent times, of attempts to recover a low dimensional structure, namely a manifold, underlying the high dimensional data. Recovering the manifold requires the noise to be of a certain concentration and prevailing methods address this requirement by constructing an approximated manifold that is based on the tangent space estimation at each sample point. Although theoretical convergence for these methods is guaranteed, the samples are either noiseless or the noise is bounded. However, if the noise is unbounded, as is commonplace, the tangent space estimation at the noisy samples will be blurred – an undesirable outcome since fitting a manifold from the blurred tangent space might be more greatly compromised in terms of its accuracy. In this paper, we introduce a new manifold-fitting method, whereby the output manifold is constructed by directly estimating the tangent spaces at the projected points on the latent manifold, rather than at the sample points, thus reducing the error caused by the noise. Assuming the noise is unbounded, our new method has a high probability of achieving theoretical convergence, in terms of the upper bound of the distance between the estimated and latent manifold. The smoothness of the estimated manifold is also evaluated by bounding the supremum of twice difference above. Numerical simulations are conducted as part of this new method to help validate our theoretical findings and demonstrate the advantages of our method over other relevant manifold fitting methods. Finally, our method is applied to real data examples.
</description>
</item>

<item>
<title>
Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play
</title>
<link>
http://jmlr.org/papers/v26/24-1503.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1503/24-1503.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Zelai Xu, Chao Yu, Yancheng Liang, Yi Wu, Yu Wang</author>
<description>
Self-play (SP) is a popular multi-agent reinforcement learning framework for competitive games. Despite the empirical success, the theoretical properties of SP are limited to two-player settings. For team competitive games where two teams of cooperative agents compete with each other, we show a counter-example where SP cannot converge to a global Nash equilibrium (NE) with high probability. Policy-Space Response Oracles (PSRO) is an alternative framework that finds NEs by iteratively learning the best response (BR) to previous policies. PSRO can be directly extended to team competitive games with unchanged convergence properties by learning team BRs, but its repeated training from scratch makes it hard to scale to complex games. In this work, we propose Generalized Fictitious Cross-Play (GFXP), a novel algorithm that inherits benefits from both frameworks. GFXP simultaneously trains an SP-based main policy and a counter population. The main policy is trained by fictitious self-play and cross-play against the counter population, while the counter policies are trained as the BRs to the main policy&#39;s checkpoints. We evaluate GFXP in matrix games and gridworld domains where GFXP achieves the lowest exploitabilities. We further conduct experiments in a challenging football game where GFXP defeats SOTA models with over 94% win rate.
</description>
</item>

<item>
<title>
Wasserstein Convergence Guarantees for a General Class of Score-Based Generative Models
</title>
<link>
http://jmlr.org/papers/v26/24-0902.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0902/24-0902.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Xuefeng Gao, Hoang M. Nguyen, Lingjiong Zhu</author>
<description>
Score-based generative models are a recent class of deep generative models with state-of-the-art performance in many applications. In this paper, we establish convergence guarantees for a general class of score-based generative models in the 2-Wasserstein distance, assuming accurate score estimates and smooth log-concave data distribution. We specialize our results to several concrete score-based generative models with specific choices of forward processes modeled by stochastic differential equations, and obtain an upper bound on the iteration complexity for each model, which demonstrates the impacts of different choices of the forward processes. We also provide a lower bound when the data distribution is Gaussian. Numerically, we experiment with score-based generative models with different forward processes for unconditional image generation on CIFAR-10. We find that the experimental results are in good agreement with our theoretical predictions on the iteration complexity.
</description>
</item>

<item>
<title>
Extremal graphical modeling with latent variables via convex optimization
</title>
<link>
http://jmlr.org/papers/v26/24-0472.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0472/24-0472.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Sebastian Engelke, Armeen Taeb</author>
<description>
Extremal graphical models encode the conditional independence structure of multivariate extremes and provide a powerful tool for quantifying the risk of rare events. Prior work on learning these graphs from data has focused on the setting where all relevant variables are observed. For the popular class of Husler-Reiss models, we propose the eglatent method, a tractable convex program for learning extremal graphical models in the presence of latent variables. Our approach decomposes the Husler-Reiss precision matrix into a sparse component encoding the graphical structure among the observed variables after conditioning on the latent variables, and a low-rank component encoding the effect of a few latent variables on the observed variables. We provide finite-sample guarantees of eglatent and show that it consistently recovers the conditional graph as well as the number of latent variables. We highlight the improved performances of our approach on synthetic and real data.
</description>
</item>

<item>
<title>
On the Approximation of Kernel functions
</title>
<link>
http://jmlr.org/papers/v26/24-0270.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0270/24-0270.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Paul Dommel, Alois Pichler</author>
<description>
Various methods in statistical learning build on kernels considered in reproducing kernel Hilbert spaces. In applications, the kernel is often selected based on characteristics of the problem and the data. This kernel is then employed to infer response variables at points, where no explanatory data were observed. The data considered here are located in compact sets in higher dimensions and the paper addresses approximations of the kernel itself. The new approach considers Taylor series approximations of radial kernel functions. For the Gauss kernel on the unit cube, the paper establishes an upper bound of the associated eigenfunctions, which grows only polynomially with respect to the index. The novel approach substantiates smaller regularization parameters than considered in the literature, overall leading to better approximations. This improvement confirms low rank approximation methods such as the Nyström method.
</description>
</item>

<item>
<title>
Efficient and Robust Semi-supervised Estimation of Average Treatment Effect with Partially Annotated Treatment and Response
</title>
<link>
http://jmlr.org/papers/v26/23-1587.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1587/23-1587.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jue Hou, Rajarshi Mukherjee, Tianxi Cai</author>
<description>
A notable challenge of leveraging Electronic Health Records (EHR) for treatment effect assessment is the lack of precise information on important clinical variables, including the treatment received and the response. Both treatment information and response cannot be accurately captured by readily available EHR features in many studies and require labor-intensive manual chart review to precisely annotate, which limits the number of available gold standard labels on these key variables. We considered average treatment effect (ATE) estimation when 1) exact treatment and outcome variables are only observed together in a small labeled subset and 2) noisy surrogates of treatment and outcome, such as relevant prescription and diagnosis codes, along with potential confounders are observed for all subjects. We derived the efficient influence function for ATE and used it to construct a semi-supervised multiple machine learning (SMMAL) estimator. We justified that our SMMAL ATE estimator is semi-parametric efficient with B-spline regression under low-dimensional smooth models. We developed the adaptive sparsity/model doubly robust estimation under high-dimensional logistic propensity score and outcome regression models. Results from simulation studies demonstrated the validity of our SMMAL method and its superiority over supervised and unsupervised benchmarks. We applied SMMAL to the assessment of targeted therapies for metastatic colorectal cancer in comparison to chemotherapy.
</description>
</item>

<item>
<title>
Nonconvex Stochastic Bregman Proximal Gradient Method with Application to Deep Learning
</title>
<link>
http://jmlr.org/papers/v26/23-0657.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0657/23-0657.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Kuangyu Ding, Jingyang Li, Kim-Chuan Toh</author>
<description>
Stochastic gradient methods for minimizing nonconvex composite objective functions typically rely on the Lipschitz smoothness of the differentiable part, but this assumption fails in many important problem classes like quadratic inverse problems and neural network training, leading to instability of the algorithms in both theory and practice. To address this, we propose a family of stochastic Bregman proximal gradient (SBPG) methods that only require smooth adaptivity. SBPG replaces the quadratic approximation in SGD with a Bregman proximity measure, offering a better approximation model that handles non-Lipschitz gradients in nonconvex objectives. We establish the convergence properties of vanilla SBPG and show it achieves optimal sample complexity in the nonconvex setting. Experimental results on quadratic inverse problems demonstrate SBPG&#39;s robustness in terms of stepsize selection and sensitivity to the initial point. Furthermore, we introduce a momentum-based variant, MSBPG, which enhances convergence by relaxing the mini-batch size requirement while preserving the optimal oracle complexity. We apply MSBPG to the training of deep neural networks, utilizing a polynomial kernel function to ensure smooth adaptivity of the loss function. Experimental results on benchmark datasets confirm the effectiveness and robustness of MSBPG in training neural networks. Given its negligible additional computational cost compared to SGD in large-scale optimization, MSBPG shows promise as a universal open-source optimizer for future applications.
</description>
</item>

<item>
<title>
Optimizing Data Collection for Machine Learning
</title>
<link>
http://jmlr.org/papers/v26/23-0292.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0292/23-0292.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Rafid Mahmood, James Lucas, Jose M. Alvarez, Sanja Fidler, Marc T. Law</author>
<description>
Modern deep learning systems require huge data sets to achieve impressive performance, but there is little guidance on how much or what kind of data to collect. Over-collecting data incurs unnecessary present costs, while under-collecting may incur future costs and delay workflows. We propose a new paradigm to model the data collection workflow as a formal optimal data collection problem that allows designers to specify performance targets, collection costs, a time horizon, and penalties for failing to meet the targets. This formulation generalizes to tasks with multiple data sources, such as labeled and unlabeled data used in semi-supervised learning, and can be easily modified to customized analyses such as how to introduce data from new classes to an existing model. To solve our problem, we develop Learn-Optimize-Collect (LOC), which minimizes expected future collection costs. Finally, we numerically compare our framework to the conventional baseline of estimating data requirements by extrapolating from neural scaling laws. We significantly reduce the risks of failing to meet desired performance targets on several classification, segmentation, and detection tasks, while maintaining low total collection costs.
</description>
</item>

<item>
<title>
Unbalanced Kantorovich-Rubinstein distance, plan, and barycenter on nite spaces: A statistical perspective
</title>
<link>
http://jmlr.org/papers/v26/22-1262.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-1262/22-1262.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Shayan Hundrieser, Florian Heinemann, Marcel Klatt, Marina Struleva, Axel Munk</author>
<description>
We analyze statistical properties of plug-in estimators for unbalanced optimal transport quantities between finitely supported measures in different prototypical sampling models. Specifically, our main results provide non-asymptotic bounds on the expected error of empirical Kantorovich-Rubinstein (KR) distance, plans, and barycenters for mass penalty parameter $C&gt;0$. The impact of the mass penalty parameter $C$ is studied in detail. Based on this analysis, we mathematically justify randomized computational schemes for KR quantities which can be used for fast approximate computations in combination with any exact solver. Using synthetic and real datasets, we empirically analyze the behavior of the expected errors in simulation studies and illustrate the validity of our theoretical bounds.
</description>
</item>

<item>
<title>
Copula-based Sensitivity Analysis for Multi-Treatment Causal Inference with Unobserved Confounding
</title>
<link>
http://jmlr.org/papers/v26/22-0372.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0372/22-0372.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jiajing Zheng, Alexander D&#39;Amour, Alexander Franks</author>
<description>
Recent work has focused on the potential and pitfalls of causal identification in observational studies with multiple simultaneous treatments. Building on previous work, we show that even if the conditional distribution of unmeasured confounders given treatments were known exactly, the causal effects would not in general be identifiable, although they may be partially identified.  Given these results, we propose a sensitivity analysis method for characterizing the effects of potential unmeasured confounding, tailored to the multiple treatment setting, that can be used to characterize a range of causal effects that are compatible with the observed data. Our method is based on a copula factorization of the joint distribution of outcomes, treatments, and confounders, and can be layered on top of arbitrary observed data models. We propose a practical implementation of this approach making use of the Gaussian copula, and establish conditions under which causal effects can be bounded. We also describe approaches for reasoning about effects, including calibrating sensitivity parameters, quantifying robustness of effect estimates, and selecting models that are most consistent with prior hypotheses.
</description>
</item>

<item>
<title>
Rank-one Convexification for Sparse Regression
</title>
<link>
http://jmlr.org/papers/v26/19-159.html
</link>
<pdf>
http://jmlr.org/papers/volume26/19-159/19-159.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Alper Atamturk, Andres Gomez</author>
<description>
Sparse regression models are increasingly prevalent due to their ease of interpretability and superior out-of-sample performance. However, the exact model of sparse regression with an $\ell_0$-constraint restricting the support of the estimators is a challenging (\NP-hard) non-convex optimization problem. In this paper, we derive new strong convex relaxations for sparse regression. These relaxations are based on the convex-hull formulations for rank-one quadratic terms with indicator variables. The new relaxations can be formulated as semidefinite optimization problems in an extended space and are stronger and more general than the state-of-the-art formulations, including the perspective reformulation and formulations with the reverse Huber penalty and the minimax concave penalty functions. Furthermore, the proposed rank-one strengthening can be interpreted as a non-separable, non-convex, unbiased sparsity-inducing regularizer, which dynamically adjusts its penalty according to the shape of the error function without inducing bias for the sparse solutions. In our computational experiments with benchmark datasets, the proposed conic formulations are solved within seconds and result in near-optimal solutions (with 0.4\% optimality gap on average) for non-convex $\ell_0$-problems. Moreover, the resulting estimators also outperform alternative convex approaches, such as lasso and elastic net regression, from a statistical perspective, achieving high prediction accuracy and good interpretability.
</description>
</item>

<item>
<title>
gsplat: An Open-Source Library for Gaussian Splatting
</title>
<link>
http://jmlr.org/papers/v26/24-1476.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1476/24-1476.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Vickie Ye, Ruilong Li, Justin Kerr, Matias Turkulainen, Brent Yi, Zhuoyang Pan, Otto Seiskari, Jianbo Ye, Jeffrey Hu, Matthew Tancik, Angjoo Kanazawa</author>
<description>
gsplat is an open-source library designed for training and developing Gaussian Splatting methods. It features a front-end with Python bindings compatible with the PyTorch library and a back-end with highly optimized CUDA kernels. gsplat offers numerous features that enhance the optimization of Gaussian Splatting models, which include optimization improvements for speed, memory, and convergence times. Experimental results demonstrate that gsplat achieves up to 10% less training time and 4x less memory than the original implementation. Utilized in several research projects, gsplat is actively maintained on GitHub. Source code is available at https://github.com/nerfstudio-project/gsplat under Apache License 2.0. We welcome contributions from the open-source community.
</description>
</item>

<item>
<title>
Statistical Inference of Constrained Stochastic Optimization via Sketched Sequential Quadratic Programming
</title>
<link>
http://jmlr.org/papers/v26/24-0530.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0530/24-0530.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Sen Na, Michael Mahoney</author>
<description>
We consider online statistical inference of constrained stochastic nonlinear optimization problems. We apply the Stochastic Sequential Quadratic Programming (StoSQP) method to solve these problems, which can be regarded as applying second-order Newton&#39;s method to the Karush-Kuhn-Tucker (KKT) conditions. In each iteration, the StoSQP method computes the Newton direction by solving a quadratic program, and then selects a proper adaptive stepsize $\bar{\alpha}_t$ to update the primal-dual iterate. To reduce dominant computational cost of the method, we inexactly solve the quadratic program in each iteration by employing an iterative sketching solver. Notably, the approximation error of the sketching solver need not vanish as iterations proceed, meaning that the per-iteration computational cost does not blow up. For the above StoSQP method, we show that under mild assumptions, the rescaled primal-dual sequence $1/\sqrt{\bar{\alpha}_t}\cdot (x_t -x^\star, \lambda_t - \lambda^\star)$ converges to a mean-zero Gaussian distribution with a nontrivial covariance matrix depending on the underlying sketching distribution. To perform inference in practice, we also analyze a plug-in covariance matrix estimator. We illustrate the asymptotic normality result of the method both on benchmark nonlinear problems in CUTEst test set and on linearly/nonlinearly constrained regression problems.
</description>
</item>

<item>
<title>
Sliced-Wasserstein Distances and Flows on Cartan-Hadamard Manifolds
</title>
<link>
http://jmlr.org/papers/v26/24-0359.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0359/24-0359.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Clément Bonet, Lucas Drumetz, Nicolas Courty</author>
<description>
While many Machine Learning methods have been developed or transposed on Riemannian manifolds to tackle data with known non-Euclidean geometry, Optimal Transport (OT) methods on such spaces have not received much attention. The main OT tool on these spaces is the Wasserstein distance, which suffers from a heavy computational burden. On Euclidean spaces, a popular alternative is the Sliced-Wasserstein distance, which leverages a closed-form solution of the Wasserstein distance in one dimension, but which is not readily available on manifolds. In this work, we derive general constructions of Sliced-Wasserstein distances on Cartan-Hadamard manifolds, Riemannian manifolds with non-positive curvature, which include among others Hyperbolic spaces or the space of Symmetric Positive Definite matrices. Then, we propose different applications such as classification of documents with a suitably learned ground cost on a manifold, and data set comparison on a product manifold. Additionally, we derive non-parametric schemes to minimize these new distances by approximating their Wasserstein gradient flows.
</description>
</item>

<item>
<title>
Accelerating optimization over the space of probability measures
</title>
<link>
http://jmlr.org/papers/v26/23-1288.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1288/23-1288.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Shi Chen, Qin Li, Oliver Tse, Stephen J. Wright</author>
<description>
The acceleration of gradient-based optimization methods is a subject of significant practical and theoretical importance, particularly within machine learning applications. While much attention has been directed towards optimizing within Euclidean space, the need to optimize over spaces of probability measures in machine learning motivates the exploration of accelerated gradient methods in this context, too. To this end, we introduce a Hamiltonian-flow approach analogous to momentum-based approaches in Euclidean space. We demonstrate that, in the continuous-time setting, algorithms based on this approach can achieve convergence rates of arbitrarily high order. We complement our findings with numerical examples.
</description>
</item>

<item>
<title>
Bayesian Multi-Group Gaussian Process Models for Heterogeneous Group-Structured Data
</title>
<link>
http://jmlr.org/papers/v26/23-0291.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0291/23-0291.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Didong Li, Andrew Jones, Sudipto Banerjee, Barbara E. Engelhardt</author>
<description>
Gaussian processes are pervasive in functional data analysis, machine learning, and spatial statistics for modeling complex dependencies. Scientific data are often heterogeneous in their inputs and contain multiple known discrete groups of samples; thus, it is desirable to leverage the similarity among groups while accounting for heterogeneity across groups. We propose multi-group Gaussian processes (MGGPs) defined over $\mathbb{R}^p\times \mathscr{C}$, where $\mathscr{C}$ is a finite set representing the group label, by developing general classes of valid (positive definite) covariance functions on such domains. MGGPs are able to accurately recover relationships between the groups and efficiently share strength across samples from all groups during inference, while capturing distinct group-specific behaviors in the conditional posterior distributions. We demonstrate inference in MGGPs through simulation experiments, and we apply our proposed MGGP regression framework to gene expression data to illustrate the behavior and enhanced inferential capabilities of multi-group Gaussian processes by jointly modeling continuous and categorical variables.
</description>
</item>

<item>
<title>
Orthogonal Bases for Equivariant Graph Learning with Provable k-WL Expressive Power
</title>
<link>
http://jmlr.org/papers/v26/23-0178.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0178/23-0178.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jia He, Maggie Cheng</author>
<description>
Graph neural network (GNN) models have been widely used for learning graph-structured data. Due to the permutation-invariant requirement of graph learning tasks, a basic element in graph neural networks is the invariant and equivariant linear layers. Previous work (Maron et al., 2019b) provided a maximal collection of invariant and equivariant linear layers and a simple deep neural network model, called k-IGN, for graph data defined on k-tuples of nodes. It is shown that the expressive power of k-IGN is at least as good as the  k-Weisfeiler-Leman (WL) algorithm in graph isomorphism tests. However, the dimension of the invariant layer and equivariant layer is the k-th and 2k-th bell numbers, respectively. Such high complexity makes it computationally infeasible for k-IGNs with k &gt;= 3. In this paper, we show that a much smaller dimension for the linear layers is sufficient to achieve the same expressive power. We provide two sets of orthogonal bases for the linear layers, each with only 3(2^k-1)-k basis elements. Based on these linear layers, we develop neural network models GNN-a and GNN-b and show that for the graph data defined on k-tuples of data, GNN-a and GNN-b achieve the expressive power of the k-WL algorithm and the (k+1)-WL algorithm in graph isomorphism tests, respectively. In molecular prediction tasks on benchmark datasets, we demonstrate that low-order neural network models consisting of the proposed linear layers achieve better performance than other neural network models. In particular, order-2 GNN-b and order-3 GNN-a both have 3-WL expressive power, but use a much smaller basis and hence much less computation time than known neural network models.
</description>
</item>

<item>
<title>
Optimal Experiment Design for Causal Effect Identification
</title>
<link>
http://jmlr.org/papers/v26/22-1516.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-1516/22-1516.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Sina Akbari, Jalal Etesami, Negar Kiyavash</author>
<description>
Pearl’s do calculus is a complete axiomatic approach to learn the identifiable causal effects from observational data. When such an effect is not identifiable, it is necessary to perform a collection of often costly interventions in the system to learn the causal effect. In this work, we consider the problem of designing a collection of interventions with the minimum cost to identify the desired effect. First, we prove that this problem is NP-complete and subsequently propose an algorithm that can either find the optimal solution or a logarithmic-factor approximation of it. This is done by establishing a connection between our problem and the minimum hitting set problem. Additionally, we propose several polynomial time heuristic algorithms to tackle the computational complexity of the problem. Although these algorithms could potentially stumble on sub-optimal solutions, our simulations show that they achieve small regrets on random graphs.
</description>
</item>

<item>
<title>
Mean Aggregator is More Robust than Robust Aggregators under Label Poisoning Attacks on Distributed Heterogeneous Data
</title>
<link>
http://jmlr.org/papers/v26/24-1307.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1307/24-1307.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jie Peng, Weiyu Li, Stefan Vlaski, Qing Ling</author>
<description>
Robustness to malicious attacks is of paramount importance for distributed learning. Existing works usually consider the classical Byzantine attacks model, which assumes that some workers can send arbitrarily malicious messages to the server and disturb the aggregation steps of the distributed learning process. To defend against such worst-case Byzantine attacks, various robust aggregators have been proposed. They are proven to be effective and much superior to the often-used mean aggregator. In this paper, however, we demonstrate that the robust aggregators are too conservative for a class of weak but practical malicious attacks, known as label poisoning attacks, where the sample labels of some workers are poisoned. Surprisingly, we are able to show that the mean aggregator is more robust than the state-of-the-art robust aggregators in theory, given that the distributed data are sufficiently heterogeneous. In fact, the learning error of the mean aggregator is proven to be order-optimal in this case. Experimental results corroborate our theoretical findings, showing the superiority of the mean aggregator under label poisoning attacks.
</description>
</item>

<item>
<title>
The Blessing of Heterogeneity in Federated Q-Learning: Linear Speedup and Beyond
</title>
<link>
http://jmlr.org/papers/v26/24-0579.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0579/24-0579.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jiin Woo, Gauri Joshi, Yuejie Chi</author>
<description>
In this paper, we consider federated Q-learning, which aims to learn an optimal Q-function by periodically aggregating local Q-estimates trained on local data alone. Focusing on infinite-horizon tabular Markov decision processes, we provide sample complexity guarantees for both the synchronous and asynchronous variants of federated Q-learning, which exhibit a linear speedup with respect to the number of agents and near-optimal dependencies on other salient problem parameters. In the asynchronous setting, existing analyses of federated Q-learning, which adopt an equally weighted averaging of local Q-estimates, require that every agent covers the entire state-action space. In contrast, our improved sample complexity scales inverse proportionally to the minimum entry of the average stationary state-action occupancy distribution of all agents, thus only requiring the agents to collectively cover the entire state-action space, unveiling the blessing of heterogeneity. However, its sample complexity still suffers when the local trajectories are highly heterogeneous. In response, we propose a novel federated Q-learning algorithm with importance averaging, giving larger weights to more frequently visited state-action pairs, which achieves a robust linear speedup as if all trajectories are centrally processed, regardless of the heterogeneity of local behavior policies.
</description>
</item>

<item>
<title>
depyf: Open the Opaque Box of PyTorch Compiler for Machine Learning Researchers
</title>
<link>
http://jmlr.org/papers/v26/24-0383.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0383/24-0383.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Kaichao You, Runsheng Bai, Meng Cao, Jianmin Wang, Ion Stoica, Mingsheng Long</author>
<description>
PyTorch 2.x introduces a compiler designed to accelerate deep learning programs. However, for machine learning researchers, fully leveraging the PyTorch compiler can be challenging due to its operation at the Python bytecode level, making it appear as an opaque box. To address this, we introduce depyf, a tool designed to demystify the inner workings of the PyTorch compiler. depyf decompiles the bytecode generated by PyTorch back into equivalent source code and establishes connections between the code objects in the memory and their counterparts in source code format on the disk. This feature enables users to step through the source code line by line using debuggers, thus enhancing their understanding of the underlying processes. Notably, depyf is non-intrusive and user-friendly, primarily relying on two convenient context managers for its core functionality. The project is openly available at https://github.com/thuml/depyf and is recognized as a PyTorch ecosystem project at https://pytorch.org/blog/introducing-depyf.
</description>
</item>

<item>
<title>
The ODE Method for Stochastic Approximation and Reinforcement Learning with Markovian Noise
</title>
<link>
http://jmlr.org/papers/v26/24-0100.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0100/24-0100.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Shuze Daniel Liu, Shuhang Chen, Shangtong Zhang</author>
<description>
Stochastic approximation is a class of algorithms that update a vector iteratively, incrementally, and stochastically, including, e.g., stochastic gradient descent and temporal difference learning. One fundamental challenge in analyzing a stochastic approximation algorithm is to establish its stability, i.e., to show that the stochastic vector iterates are bounded almost surely. In this paper, we extend the celebrated Borkar-Meyn theorem for stability from the Martingale difference noise setting to the Markovian noise setting, which greatly improves its applicability in reinforcement learning, especially in those off-policy reinforcement learning algorithms with linear function approximation and eligibility traces. Central to our analysis is the diminishing asymptotic rate of change of a few functions, which is implied by both a form of the strong law of large numbers and a form of the law of the iterated logarithm.
</description>
</item>

<item>
<title>
Improving Graph Neural Networks on Multi-node Tasks with the Labeling Trick
</title>
<link>
http://jmlr.org/papers/v26/23-0560.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0560/23-0560.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Xiyuan Wang, Pan Li, Muhan Zhang</author>
<description>
In this paper, we study using graph neural networks (GNNs) for multi-node representation learning, where a representation for a set of more than one node (such as a link) is to be learned. Existing GNNs are mainly designed to learn single-node representations. When used for multi-node representation learning, a common practice is to directly aggregate the single-node representations obtained by a GNN. In this paper, we show a fundamental limitation of such an approach, namely the inability to capture the dependence among multiple nodes in the node set. A straightforward solution is to distinguish target nodes from others. Formalizing this idea, we propose \text{labeling trick}, which first labels nodes in the graph according to their relationships with the target node set before applying a GNN and then aggregates node representations obtained in the labeled graph for multi-node representations. Besides node sets in graphs, we also extend labeling tricks to posets, subsets and hypergraphs. Experiments verify that the labeling trick technique can boost GNNs on various tasks, including undirected link prediction, directed link prediction, hyperedge prediction, and subgraph prediction. Our work explains the superior performance of previous node-labeling-based methods and establishes a theoretical foundation for using GNNs for multi-node representation learning.
</description>
</item>

<item>
<title>
Directed Cyclic Graphs for Simultaneous Discovery of Time-Lagged and Instantaneous Causality from Longitudinal Data Using Instrumental Variables
</title>
<link>
http://jmlr.org/papers/v26/23-0272.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0272/23-0272.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Wei Jin, Yang Ni, Amanda B. Spence, Leah H. Rubin, Yanxun Xu</author>
<description>
We consider the problem of causal discovery from longitudinal observational data. We develop a novel framework that simultaneously discovers the time-lagged causality and the possibly cyclic instantaneous causality. Under common causal discovery assumptions, combined with additional instrumental information typically available in longitudinal data, we prove the proposed model is generally identifiable. To the best of our knowledge, this is the first causal identification theory for directed graphs with general cyclic patterns that achieves unique causal identifiability. Structural learning is carried out in a fully Bayesian fashion. Through extensive simulations and an application to the Women&#39;s Interagency HIV Study, we demonstrate the identifiability, utility, and superiority of the proposed model against state-of-the-art alternative methods.
</description>
</item>

<item>
<title>
Bayesian Sparse Gaussian Mixture Model for Clustering in High Dimensions
</title>
<link>
http://jmlr.org/papers/v26/23-0142.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0142/23-0142.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Dapeng Yao, Fangzheng Xie, Yanxun Xu</author>
<description>
We study the sparse high-dimensional Gaussian mixture model when the number of clusters is allowed to grow with the sample size. A minimax lower bound for parameter estimation is established, and we show that a constrained maximum likelihood estimator achieves the minimax lower bound. However, this optimization-based estimator is computationally intractable because the objective function is highly nonconvex and the feasible set involves discrete structures. To address the computational challenge, we propose a computationally tractable Bayesian approach to estimate high-dimensional Gaussian mixtures whose cluster centers exhibit sparsity using a continuous spike-and-slab prior. We further prove that the posterior contraction rate of the proposed Bayesian method is minimax optimal. The mis- clustering rate is obtained as a by-product using tools from matrix perturbation theory. The proposed Bayesian sparse Gaussian mixture model does not require pre-specifying the number of clusters, which can be adaptively estimated. The validity and usefulness of the proposed method is demonstrated through simulation studies and the analysis of a real-world single-cell RNA sequencing data set.
</description>
</item>

<item>
<title>
Regularizing Hard Examples Improves Adversarial Robustness
</title>
<link>
http://jmlr.org/papers/v26/22-1428.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-1428/22-1428.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Hyungyu Lee, Saehyung Lee, Ho Bae, Sungroh Yoon</author>
<description>
Recent studies have validated that pruning hard-to-learn examples from training improves the generalization performance of neural networks (NNs). In this study, we investigate this intriguing phenomenon---the negative effect of hard examples on generalization---in adversarial training. Particularly, we theoretically demonstrate that the increase in the difficulty of hard examples in adversarial training is significantly greater than the increase in the difficulty of easy examples. Furthermore, we verify that hard examples are only fitted through memorization of the label in adversarial training. We conduct both theoretical and empirical analyses of this memorization phenomenon, showing that pruning hard examples in adversarial training can enhance the model&#39;s robustness. However, the challenge remains in finding the optimal threshold for removing hard examples that degrade robustness performance. Based upon these observations, we propose a new approach, difficulty proportional label smoothing (DPLS), to adaptively mitigate the negative effect of hard examples, thereby improving the adversarial robustness of NNs. Notably, our experimental result indicates that our method can successfully leverage hard examples while circumventing the negative effect.
</description>
</item>

<item>
<title>
Random ReLU Neural Networks as Non-Gaussian Processes
</title>
<link>
http://jmlr.org/papers/v26/24-0737.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0737/24-0737.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Rahul Parhi, Pakshal Bohra, Ayoub El Biari, Mehrsa Pourya, Michael Unser</author>
<description>
We consider a large class of shallow neural networks with randomly initialized parameters and rectified linear unit activation functions. We prove that these random neural networks are well-defined non-Gaussian processes. As a by-product, we demonstrate that these networks are solutions to stochastic differential equations driven by impulsive white noise (combinations of random Dirac measures). These processes are parameterized by the law of the weights and biases as well as the density of activation thresholds in each bounded region of the input domain. We prove that these processes are isotropic and wide-sense self-similar with Hurst exponent 3/2. We also derive a remarkably simple closed-form expression for their autocovariance function. Our results are fundamentally different from prior work in that we consider a non-asymptotic viewpoint: The number of neurons in each bounded region of the input domain (i.e., the width) is itself a random variable with a Poisson law with mean proportional to the density parameter. Finally, we show that, under suitable hypotheses, as the expected width tends to infinity, these processes can converge in law not only to Gaussian processes, but also to non-Gaussian processes depending on the law of the weights. Our asymptotic results provide a new take on several classical results (wide networks converge to Gaussian processes) as well as some new ones (wide networks can converge to non-Gaussian processes).
</description>
</item>

<item>
<title>
Riemannian Bilevel Optimization
</title>
<link>
http://jmlr.org/papers/v26/24-0397.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0397/24-0397.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jiaxiang Li, Shiqian Ma</author>
<description>
In this work, we consider the bilevel optimization problem on Riemannian manifolds. We inspect the calculation of the hypergradient of such problems on general manifolds and thus enable the utilization of gradient-based algorithms to solve such problems. The calculation of the hypergradient requires utilizing the notion of Riemannian cross-derivative and we inspect the properties and the numerical calculations of Riemannian cross-derivatives. Algorithms in both deterministic and stochastic settings, named respectively RieBO and RieSBO, are proposed that include the existing Euclidean bilevel optimization algorithms as special cases. Numerical experiments on robust optimization on Riemannian manifolds are presented to show the applicability and efficiency of the proposed methods.
</description>
</item>

<item>
<title>
Supervised Learning with Evolving Tasks and Performance Guarantees
</title>
<link>
http://jmlr.org/papers/v26/24-0343.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0343/24-0343.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Verónica Álvarez, Santiago Mazuelas, Jose A. Lozano</author>
<description>
Multiple supervised learning scenarios are composed by a sequence of classification tasks. For instance, multi-task learning and continual learning aim to learn a sequence of tasks that is either fixed or grows over time. Existing techniques for learning tasks that are in a sequence are tailored to specific scenarios, lacking adaptability to others. In addition, most of existing techniques consider situations in which the order of the tasks in the sequence is not relevant. However, it is common that tasks in a sequence are evolving in the sense that consecutive tasks often have a higher similarity. This paper presents a learning methodology that is applicable to multiple supervised learning scenarios and adapts to evolving tasks. Differently from existing techniques, we provide computable tight performance guarantees and analytically characterize the increase in the effective sample size. Experiments on benchmark datasets show the performance improvement of the proposed methodology in multiple scenarios and the reliability of the presented performance guarantees.
</description>
</item>

<item>
<title>
Error estimation and adaptive tuning for unregularized robust M-estimator
</title>
<link>
http://jmlr.org/papers/v26/24-0060.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0060/24-0060.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Pierre C. Bellec, Takuya Koriyama</author>
<description>
We consider unregularized robust M-estimators for linear models under Gaussian design and heavy-tailed noise, in the proportional asymptotics regime where the sample size n and the number of features p are both increasing such that $p/n \to \gamma\in (0,1)$. An estimator of the out-of-sample error of a robust M-estimator is analyzed and proved to be consistent for a large family of loss functions that includes the Huber loss. As an application of this result, we propose an adaptive tuning procedure of the scale parameter $\lambda&gt;0$ of a given loss function $\rho$: choosing $\hat \lambda$ in a given interval $I$ that minimizes the out-of-sample error estimate of the M-estimator constructed with loss $\rho_\lambda(\cdot) = \lambda^2 \rho(\cdot/\lambda)$ leads to the optimal out-of-sample error over $I$. The proof relies on a smoothing argument: the unregularized M-estimation objective function is perturbed, or smoothed, with a Ridge penalty that vanishes as $n\to+\infty$, and shows that the unregularized M-estimator of interest inherits properties of its smoothed version.
</description>
</item>

<item>
<title>
From Sparse to Dense Functional Data in High Dimensions: Revisiting Phase Transitions from a Non-Asymptotic Perspective
</title>
<link>
http://jmlr.org/papers/v26/23-1578.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1578/23-1578.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Shaojun Guo, Dong Li, Xinghao Qiao, Yizhu Wang</author>
<description>
Nonparametric estimation of the mean and covariance functions is ubiquitous in functional data analysis and local linear smoothing techniques are most frequently used. Zhang and Wang (2016) explored different types of asymptotic properties of the estimation, which reveal interesting phase transition phenomena based on the relative order of the average sampling frequency per subject $T$ to the number of subjects $n$, partitioning the data into three categories: “sparse”, “semi-dense”, and “ultra-dense”. In an increasingly available high-dimensional scenario, where the number of functional variables $p$ is large in relation to $n$, we revisit this open problem from a non-asymptotic perspective by deriving comprehensive concentration inequalities for the local linear smoothers. Besides being of interest by themselves, our non-asymptotic results lead to elementwise maximum rates of $L_2$ convergence and uniform convergence serving as a fundamentally important tool for further convergence analysis when $p$ grows exponentially with $n$ and possibly $T$. With the presence of extra $\log p$ terms to account for the high-dimensional effect, we then investigate the scaled phase transitions and the corresponding elementwise maximum rates from sparse to semi-dense to ultra-dense functional data in high dimensions. We also discuss a couple of applications of our theoretical results. Finally, numerical studies are carried out to confirm the established theoretical properties.
</description>
</item>

<item>
<title>
Locally Private Causal Inference for Randomized Experiments
</title>
<link>
http://jmlr.org/papers/v26/23-1401.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1401/23-1401.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Yuki Ohnishi, Jordan Awan</author>
<description>
Local differential privacy is a differential privacy paradigm in which individuals first apply a privacy mechanism to their data (often by adding noise) before transmitting the result to a curator. The noise for privacy results in additional bias and variance in their analyses. Thus it is of great importance for analysts to incorporate the privacy noise into valid inference. In this article, we develop methodologies to infer causal effects from locally privatized data under randomized experiments. First, we present frequentist estimators under various privacy scenarios with their variance estimators and plug-in confidence intervals. We show a na\&#34;ive debiased estimator results in inferior mean-squared error (MSE) compared to minimax lower bounds. In contrast, we show that using a customized privacy mechanism, we can match the lower bound, giving minimax optimal inference. We also develop a Bayesian nonparametric methodology along with a blocked Gibbs sampling algorithm, which can be applied to any of our proposed privacy mechanisms, and which performs especially well in terms of MSE for tight privacy budgets. Finally, we present simulation studies to evaluate the performance of our proposed frequentist and Bayesian methodologies for various privacy budgets, resulting in useful suggestions for performing causal inference for privatized data.
</description>
</item>

<item>
<title>
Estimating Network-Mediated Causal Effects via Principal Components Network Regression
</title>
<link>
http://jmlr.org/papers/v26/23-1317.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1317/23-1317.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Alex Hayes, Mark M. Fredrickson, Keith Levin</author>
<description>
We develop a method to decompose causal effects on a social network into an indirect effect mediated by the network, and a direct effect independent of the social network. To handle the complexity of network structures, we assume that latent social groups act as causal mediators. We develop principal components network regression models to differentiate the social effect from the non-social effect. Fitting the regression models is as simple as principal components analysis followed by ordinary least squares estimation. We prove asymptotic theory for regression coefficients from this procedure and show that it is widely applicable, allowing for a variety of distributions on the regression errors and network edges. We carefully characterize the counterfactual assumptions necessary to use the regression models for causal inference, and show that current approaches to causal network regression may result in over-control bias. The method is very general, so that it is applicable to many types of structured data beyond social networks, such as text, areal data, psychometrics, images and omics.
</description>
</item>

<item>
<title>
Selective Inference with Distributed Data
</title>
<link>
http://jmlr.org/papers/v26/23-0309.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-0309/23-0309.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Sifan Liu, Snigdha Panigrahi</author>
<description>
When data are distributed across multiple sites or machines rather than centralized in one location, researchers face the challenge of extracting meaningful information without directly sharing individual data points. While there are many distributed methods for point estimation using sparse regression, few options are available for estimating uncertainties or conducting hypothesis tests based on the estimated sparsity. In this paper, we introduce a procedure for performing selective inference with distributed data. We consider a scenario where each local machine solves a lasso problem and communicates the selected predictors to a central machine. The central machine then aggregates these selected predictors to form a generalized linear model (GLM). Our goal is to provide valid inference for the selected GLM while reusing data that have been used in the model selection process. Our proposed procedure only requires low-dimensional summary statistics from local machines, thus keeping communication costs low and preserving the privacy of individual data sets. Furthermore, this procedure can be applied in scenarios where model selection is repeatedly conducted on randomly subsampled data sets, addressing the p-value lottery problem linked with model selection. We demonstrate the effectiveness of our approach through simulations and an analysis of a medical data set on ICU admissions.
</description>
</item>

<item>
<title>
Two-Timescale Gradient Descent Ascent Algorithms for Nonconvex Minimax Optimization
</title>
<link>
http://jmlr.org/papers/v26/22-0863.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0863/22-0863.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Tianyi Lin, Chi Jin, Michael I. Jordan</author>
<description>
We provide a unified analysis of two-timescale gradient descent ascent (TTGDA) for solving structured nonconvex minimax optimization problems in the form of $\min_x \max_{y \in Y} f(x, y)$, where the objective function $f(x, y)$ is nonconvex in $x$ and concave in $y$, and the constraint set $Y \subseteq \mathbb{R}^n$ is convex and bounded. In the convex-concave setting, the single-timescale gradient descent ascent (GDA) algorithm is widely used in applications and has been shown to have strong convergence guarantees. In more general settings, however, it can fail to converge. Our contribution is to design TTGDA algorithms that are effective beyond the convex-concave setting, efficiently finding a stationary point of the function $\Phi(\cdot) := \max_{y \in Y} f(\cdot, y)$. We also establish theoretical bounds on the complexity of solving both smooth and nonsmooth nonconvex-concave minimax optimization problems. To the best of our knowledge, this is the first systematic analysis of TTGDA for nonconvex minimax optimization, shedding light on its superior performance in training generative adversarial networks (GANs) and in other real-world application problems.
</description>
</item>

<item>
<title>
An Axiomatic Definition of Hierarchical Clustering
</title>
<link>
http://jmlr.org/papers/v26/24-1052.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-1052/24-1052.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Ery Arias-Castro, Elizabeth Coda</author>
<description>
In this paper, we take an axiomatic approach to defining a population hierarchical clustering for piecewise constant densities, and in a similar manner to Lebesgue integration, extend this definition to more general densities. When the density satisfies some mild conditions, e.g., when it has connected support, is continuous, and vanishes only at infinity, or when the connected components of the density satisfy these conditions, our axiomatic definition results in Hartigan&#39;s definition of cluster tree.
</description>
</item>

<item>
<title>
Test-Time Training on Video Streams
</title>
<link>
http://jmlr.org/papers/v26/24-0439.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0439/24-0439.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Renhao Wang, Yu Sun, Arnuv Tandon, Yossi Gandelsman, Xinlei Chen, Alexei A. Efros, Xiaolong Wang</author>
<description>
Prior work has established Test-Time Training (TTT) as a general framework to further improve a trained model at test time. Before making a prediction on each test instance, the model is first trained on the same instance using a self-supervised task such as reconstruction. We extend TTT to the streaming setting, where multiple test instances - video frames in our case - arrive in temporal order. Our extension is online TTT: The current model is initialized from the previous model, then trained on the current frame and a small window of frames immediately before. Online TTT significantly outperforms the fixed-model baseline for four tasks, on three real-world datasets. The improvements are more than 2.2x and 1.5x for instance and panoptic segmentation. Surprisingly, online TTT also outperforms its offline variant that accesses strictly more information, training on all frames from the entire test video regardless of temporal order. This finding challenges those in prior work using synthetic videos. We formalize a notion of locality as the advantage of online over offline TTT, and analyze its role with ablations and a theory based on bias-variance trade-off.
</description>
</item>

<item>
<title>
Adaptive Client Sampling in Federated Learning via Online Learning with Bandit Feedback
</title>
<link>
http://jmlr.org/papers/v26/24-0385.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0385/24-0385.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Boxin Zhao, Lingxiao Wang, Ziqi Liu, Zhiqiang Zhang, Jun Zhou, Chaochao Chen, Mladen Kolar</author>
<description>
Due to the high cost of communication, federated learning (FL) systems need to sample a subset of clients that are involved in each round of training. As a result, client sampling plays an important role in FL systems as it affects the convergence rate of optimization algorithms used to train machine learning models. Despite its importance, there is limited work on how to sample clients effectively. In this paper, we cast client sampling as an online learning task with bandit feedback, which we solve with an online stochastic mirror descent (OSMD) algorithm designed to minimize the sampling variance. We then theoretically show how our sampling method can improve the convergence speed of federated optimization algorithms over the widely used uniform sampling. Through both simulated and real data experiments, we empirically illustrate the advantages of the proposed client sampling algorithm over uniform sampling and existing online learning-based sampling strategies. The proposed adaptive sampling procedure is applicable beyond the FL problem studied here and can be used to improve the performance of stochastic optimization procedures such as stochastic gradient descent and stochastic coordinate descent.
</description>
</item>

<item>
<title>
A Random Matrix Approach to Low-Multilinear-Rank Tensor Approximation
</title>
<link>
http://jmlr.org/papers/v26/24-0193.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0193/24-0193.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Hugo Lebeau, Florent Chatelain, Romain Couillet</author>
<description>
This work presents a comprehensive understanding of the estimation of a planted low-rank signal from a general spiked tensor model near the computational threshold. Relying on standard tools from the theory of large random matrices, we characterize the large-dimensional spectral behavior of the unfoldings of the data tensor and exhibit relevant signal-to-noise ratios governing the detectability of the principal directions of the signal. These results allow to accurately predict the reconstruction performance of truncated multilinear SVD (MLSVD) in the non-trivial regime. This is particularly important since it serves as an initialization of the higher-order orthogonal iteration (HOOI) scheme, whose convergence to the best low-multilinear-rank approximation depends entirely on its initialization. We give a sufficient condition for the convergence of HOOI and show that the number of iterations before convergence tends to $1$ in the large-dimensional limit.
</description>
</item>

<item>
<title>
Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of Agents
</title>
<link>
http://jmlr.org/papers/v26/24-0043.html
</link>
<pdf>
http://jmlr.org/papers/volume26/24-0043/24-0043.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Marco Pleines, Matthias Pallasch, Frank Zimmer, Mike Preuss</author>
<description>
Memory Gym presents a suite of 2D partially observable environments, namely Mortar Mayhem, Mystery Path, and Searing Spotlights, designed to benchmark memory capabilities in decision-making agents. These environments, originally with finite tasks, are expanded into innovative, endless formats, mirroring the escalating challenges of cumulative memory games such as “I packed my bag”. This progression in task design shifts the focus from merely assessing sample efficiency to also probing the levels of memory effectiveness in dynamic, prolonged scenarios. To address the gap in available memory-based Deep Reinforcement Learning baselines, we introduce an implementation within the open-source CleanRL library that integrates Transformer-XL (TrXL) with Proximal Policy Optimization. This approach utilizes TrXL as a form of episodic memory, employing a sliding window technique. Our comparative study between the Gated Recurrent Unit (GRU) and TrXL reveals varied performances across our finite and endless tasks. TrXL, on the finite environments, demonstrates superior effectiveness over GRU, but only when utilizing an auxiliary loss to reconstruct observations. Notably, GRU makes a remarkable resurgence in all endless tasks, consistently outperforming TrXL by significant margins. Website and Source Code: https://marcometer.github.io/jmlr_2024.github.io/
</description>
</item>

<item>
<title>
Enhancing Graph Representation Learning with Localized Topological Features
</title>
<link>
http://jmlr.org/papers/v26/23-1424.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1424/23-1424.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Zuoyu Yan, Qi Zhao, Ze Ye, Tengfei Ma, Liangcai Gao, Zhi Tang, Yusu Wang, Chao Chen</author>
<description>
Representation learning on graphs is a fundamental problem that can be crucial in various tasks. Graph neural networks, the dominant approach for graph representation learning, are limited in their representation power. Therefore, it can be beneficial to explicitly extract and incorporate high-order topological and geometric information into these models. In this paper, we propose a principled approach to extract the rich connectivity information of graphs based on the theory of persistent homology. Our method utilizes the topological features to enhance the representation learning of graph neural networks and achieve state-of-the-art performance on various node classification and link prediction benchmarks. We also explore the option of end-to-end learning of the topological features, i.e., treating topological computation as a differentiable operator during learning. Our theoretical analysis and empirical study provide insights and potential guidelines for employing topological features in graph learning tasks.
</description>
</item>

<item>
<title>
Deep Out-of-Distribution Uncertainty Quantification via Weight Entropy Maximization
</title>
<link>
http://jmlr.org/papers/v26/23-1359.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1359/23-1359.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Antoine de Mathelin, François Deheeger, Mathilde Mougeot, Nicolas Vayatis</author>
<description>
This paper deals with uncertainty quantification and out-of-distribution detection in deep learning using Bayesian and ensemble methods. It proposes a practical solution to the lack of prediction diversity observed recently for standard approaches when used out-of-distribution (Ovadia et al., 2019; Liu et al., 2021). Considering that this issue is mainly related to a lack of weight diversity, we claim that standard methods sample in &#34;over-restricted&#34; regions of the weight space due to the use of &#34;over-regularization&#34; processes, such as weight decay and zero-mean centered Gaussian priors. We propose to solve the problem by adopting the maximum entropy principle for the weight distribution, with the underlying idea to maximize the weight diversity. Under this paradigm, the epistemic uncertainty is described by the weight distribution of maximal entropy that produces neural networks &#34;consistent&#34; with the training observations. Considering stochastic neural networks, a practical optimization is derived to build such a distribution, defined as a trade-off between the average empirical risk and the weight distribution entropy. We provide both theoretical and numerical results to assess the efficiency of the approach. In particular, the proposed algorithm appears in the top three best methods in all configurations of an extensive out-of-distribution detection benchmark including more than thirty competitors.
</description>
</item>

<item>
<title>
DisC2o-HD: Distributed causal inference with covariates shift for analyzing real-world high-dimensional data
</title>
<link>
http://jmlr.org/papers/v26/23-1254.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-1254/23-1254.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Jiayi Tong, Jie Hu, George Hripcsak, Yang Ning, Yong Chen</author>
<description>
High-dimensional healthcare data, such as electronic health records (EHR) data and claims data, present two primary challenges due to the large number of variables and the need to consolidate data from multiple clinical sites. The third key challenge is the potential existence of heterogeneity in terms of covariate shift. In this paper, we propose a distributed learning algorithm accounting for covariate shift to estimate the average treatment effect (ATE) for high-dimensional data, named DisC2o-HD. Leveraging the surrogate likelihood method, our method calibrates the estimates of the propensity score and outcome models to approximately attain the desired covariate balancing property, while accounting for the covariate shift across multiple clinical sites. We show that our distributed covariate balancing propensity score estimator can approximate the pooled estimator, which is obtained by pooling the data from multiple sites together. The proposed estimator remains consistent if either the propensity score model or the outcome regression model is correctly specified. The semiparametric efficiency bound is achieved when both the propensity score and the outcome models are correctly specified. We conduct simulation studies to demonstrate the performance of the proposed algorithm; additionally, we conduct an empirical study to present the readiness of implementation and validity.
</description>
</item>

<item>
<title>
Bayes Meets Bernstein at the Meta Level: an Analysis of Fast Rates in Meta-Learning with PAC-Bayes
</title>
<link>
http://jmlr.org/papers/v26/23-025.html
</link>
<pdf>
http://jmlr.org/papers/volume26/23-025/23-025.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Charles Riou, Pierre Alquier, Badr-Eddine Chérief-Abdellatif</author>
<description>
Bernstein&#39;s condition is a key assumption that guarantees fast rates in machine learning. For example, under this condition, the Gibbs posterior with prior $\pi$ has an excess risk in $O(d_{\pi}/n)$, as opposed to $O(\sqrt{d_{\pi}/n})$ in the general case, where $n$ denotes the number of observations and $d_{\pi}$ is a complexity parameter which depends on the prior $\pi$. In this paper, we examine the Gibbs posterior in the context of meta-learning, i.e., when learning the prior $\pi$ from $T$ previous tasks. Our main result is that Bernstein&#39;s condition always holds at the meta level, regardless of its validity at the observation level. This implies that the additional cost to learn the Gibbs prior $\pi$, which will reduce the term $d_\pi$ across tasks, is in $O(1/T)$, instead of the expected $O(1/\sqrt{T})$. We further illustrate how this result improves on the standard rates in three different settings: discrete priors, Gaussian priors and mixture of Gaussian priors.
</description>
</item>

<item>
<title>
Efficiently Escaping Saddle Points in Bilevel Optimization
</title>
<link>
http://jmlr.org/papers/v26/22-0136.html
</link>
<pdf>
http://jmlr.org/papers/volume26/22-0136/22-0136.pdf
</pdf>
<pubDate>2025</pubDate>
<author>Minhui Huang, Xuxing Chen, Kaiyi Ji, Shiqian Ma, Lifeng Lai</author>
<description>
Bilevel optimization is one of the fundamental problems in machine learning and optimization. Recent theoretical developments in bilevel optimization focus on finding the first-order stationary points for nonconvex-strongly-convex cases. In this paper, we analyze algorithms that can escape saddle points in nonconvex-strongly-convex bilevel optimization. Specifically, we show that the perturbed approximate implicit differentiation (AID) with a warm start strategy finds an $\epsilon$-approximate local minimum of bilevel optimization in $\tilde{O}(\epsilon^{-2})$ iterations with high probability. Moreover, we propose an inexact NEgative-curvature-Originated-from-Noise Algorithm (iNEON), an algorithm that can escape saddle point and find local minimum of stochastic bilevel optimization. As a by-product, we provide the first nonasymptotic analysis of perturbed multi-step gradient descent ascent (GDmax) algorithm that converges to local minimax point for minimax problems.
</description>
</item>


</channel>
</rss>