[{"command":"settings","settings":{"pluralDelimiter":"\u0003","suppressDeprecationErrors":true,"entitySetting":{"type":"bibcite_reference","bundle":"thesis","mapping":{"node":{"blog":"blog","class":"classes","events":"calendar","faq":"faq","link":"links","news":"news","page":"","person":"people","presentation":"presentations","software_project":"software","software_release":"software"},"bibcite_reference":{"*":"publications"},"paragraph":{"class_material":"classes"}},"viewmode":"teaser"},"mathjax":{"config_type":0,"config":{"tex2jax":{"inlineMath":[["\\(","\\)"]],"processEscapes":"true","processClass":"tex2jax_process","ignoreClass":"tex2jax_ignore","displayMath":[["$$","$$"],["\\[","\\]"]]},"showProcessingMessages":"false","messageStyle":"none"}},"user":{"uid":0,"permissionsHash":"6bef6aaaedd75fb8ec6b1c68a72fc173a14cd00edd684394ef97a6feeb8a0a42"}},"merge":true},{"command":"add_js","selector":"body","data":[{"src":"\/files\/js\/js_aI3_gL32gxGuAEBDE3LJuj3GqvgCzzvHhe76TLNNniU.js?scope=footer\u0026delta=0\u0026language=en\u0026theme=american_bold\u0026include=eJxdyMENgDAIAMCFmrKFI_g1iCRiaGkEtOP7957XMM4LJzhHjmK-jdxVCEOsOxx3DtT669rEqTwuwUDWg2ck6pKqq_D7AVDIINA"},{"src":"https:\/\/cdnjs.cloudflare.com\/ajax\/libs\/mathjax\/2.7.9\/MathJax.js?config=TeX-AMS-MML_HTMLorMML"},{"src":"\/files\/js\/js_LiGPnu81bkaTZ8S1P7qKs_weFHjgHHL9Yow35FtHTjE.js?scope=footer\u0026delta=2\u0026language=en\u0026theme=american_bold\u0026include=eJxdyMENgDAIAMCFmrKFI_g1iCRiaGkEtOP7957XMM4LJzhHjmK-jdxVCEOsOxx3DtT669rEqTwuwUDWg2ck6pKqq_D7AVDIINA"}]},{"command":"insert","method":"replaceWith","selector":"#","data":"\n\u003Cul  id=\u0022list-of-posts\u0022 more_link_id=\u0022node-readmore\u0022 class=\u0022publications view-teaser grid-view\u0022\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EPetruck, Julian. Submitted. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/epistemic-mythology-machine-learning\u0022 hreflang=\u0022en\u0022\u003EThe Epistemic Mythology of Machine Learning\u003C\/a\u003E\u201d. \u003Ci\u003EDepartment of Computer Science\u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EMuch of the epistemic legitimacy attributed to machine learning comes from a\u003Cbr\u003Efamiliar narrative about the relationship between the world, data, models, and de-\u003Cbr\u003Ecisions. According to this narrative, data objectively represent the world, models\u003Cbr\u003Ereveal its underlying patterns,and their outputs provide reliable evidence for action.\u003Cbr\u003EThis thesis argues that this story functions as an operative myth: not becausei ts in-\u003Cbr\u003Edividual steps are fictitious, but because it presents epistemic legitimacy as passing\u003Cbr\u003Eautomatically from one stage to another.\u003Cbr\u003EThe thesis follows this narrative from world to data, from data to models, and from\u003Cbr\u003Emodel outputs back into social practice. It shows how data are produced through\u003Cbr\u003Esituated practices of classification, operationalization, and quantification; how pre-\u003Cbr\u003Edictive performance establishes fit to a particular representation without by itself\u003Cbr\u003Eestablishing validity with respect to a target in the world; and how scores,rankings,\u003Cbr\u003Eand categories introduce further judgments when they are translated into interven-\u003Cbr\u003Etions. Since such interventions can also reshape the conditions under which future\u003Cbr\u003Edata are made, deployment does not end the epistemic process.\u003Cbr\u003EAn alternative account in which epistemic legitimacy is earned through warranted\u003Cbr\u003Erelations rather than inherited across a machine learning pipeline. Representations\u003Cbr\u003Eare assessed for their adequacy to particular purposes andi nferences. Objectivityi s\u003Cbr\u003Eunderstood as situated and referenced rather than perspective-free. And even well-\u003Cbr\u003Esupported predictions remain distinct from the normative justification of actions.\u003Cbr\u003EOn this account, machine learning can produce useful knowledge, but its epistemic\u003Cbr\u003Eforce depends on relations between representations,targets,claims,and actions that\u003Cbr\u003Emust remain visible,contestable, and open to revision.\u003C\/p\u003E\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EDerr, Rabanus, Alina Wernick, and Robert C. Williamson. Forthcoming. \u201c\u003Ca href=\u0022\/fmls\/publications\/law-large-numbers-accuracy-statistical-measure-ai-compliance-and-competition\u0022 hreflang=\u0022en\u0022\u003ELaw of Large Numbers: Accuracy As a Statistical Measure for AI Compliance and Competition\u003C\/a\u003E\u201d. In \u003Ci\u003EAIES2026\u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EThe machine learning community progresses (in part) by improving the\u00a0\u003Cbr\u003E\u00a0``accuracy\u0027\u0027 of its systems.\u003Cbr\u003EThe \u00a0EU AI Act explicitly refers to ``accuracy\u0027\u0027 as part of its compliance measures for high-risk AI systems.\u003Cbr\u003E\u00a0Are we talking about the same thing?\u003Cbr\u003E\u00a0This work presents ``accuracy\u0027\u0027 as a case-study for differing requirements of social worlds, the technological machine learning community and the legal community.\u003Cbr\u003E\u00a0While competition on accuracy contributes to technological development, machine learning scholars simultaneously recognize accuracy\u0027s shortcomings regarding the usefulness and effectiveness of machine learning systems.\u003Cbr\u003E\u00a0The legal counterpart embraces the vagueness of ``accuracy,\u0027\u0027 leaving interpretative flexibility for technological and societal changes. At the same time, accuracy is a core element of compliance within the EU AI Act.\u003Cbr\u003E\u00a0We elaborate on five main tensions, (a) nature of accuracy, (b) notion of performance, (c) scope of validity, (d) ends, and (e) statisticalness, to show that the two communities project disparate, and sometimes \u00a0contradictory, expectations on accuracy.\u003Cbr\u003E\u00a0Both legal and technical communities lack precise understanding of ``accuracy\u0027\u0027 beyond the contextual boundaries of their community.\u003Cbr\u003E\u00a0The resulting frictions, e.g., based on the empirical or normative understanding of accuracy, are symptoms of an unresolved (and unresolvable) debate on what accuracy is.\u003Cbr\u003E\u00a0We constructively use the frictions to recommend baselines and interventional studies in standardization, and demand for tools to extend the validity of accuracy measurements.\u003C\/p\u003E\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EDerr, Rabanus, Jessie Finocchiaro, and Robert C. Williamson. 2026. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/three-types-calibration-properties-and-theirsemantic-and-formal-relationships\u0022 hreflang=\u0022en\u0022\u003EThree Types of Calibration With Properties and TheirSemantic and Formal Relationships\u003C\/a\u003E\u201d. \u003Ci\u003EJournal of Machine Learning Research\u003C\/i\u003E 27 (110): 1-52.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/www.jmlr.org\/papers\/v27\/25-1064.html\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EFueled by discussions around \u0022trustworthiness\u0022 and algorithmic fairness, calibration of predictive systems has regained scholars attention. The vanilla definition and understanding of calibration is, simply put, on all days on which the rain probability has been predicted to be p, the actual frequency of rain days was p. However, the increased attention has led to an immense variety of new notions of \u0022calibration.\u0022 Some of the notions are incomparable, serve different purposes, or imply each other. In this work, we provide two accounts which motivate calibration: self-realization of forecasted properties and precise estimation of incurred losses of the decision makers relying on forecasts. We substantiate the former via the reflection principle and the latter by actuarial fairness. For both accounts we formulate prototypical definitions via properties $\\Gamma$ of outcome distributions, e.g., the mean or median. The prototypical definition for self-realization, which we call $\\Gamma$-calibration, is equivalent to a certain type of swap regret under certain conditions. These implications are strongly connected to the omniprediction learning paradigm. The prototypical definition for precise loss estimation is a modification of decision calibration adopted from Zhao et al. [73]. For binary outcome sets both prototypical definitions coincide under appropriate choices of reference properties. For higher-dimensional outcome sets, both prototypical definitions can be subsumed by a natural extension of the binary definition, called distribution calibration with respect to a property. We conclude by commenting on the role of groupings in both accounts of calibration often used to obtain multicalibration. In sum, this work provides a semantic map of calibration in order to navigate a fragmented terrain of notions and definitions.\u003C\/p\u003E\u003Cdiv\u003E\u003C\/div\u003E\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EIacovissi, Laura, Nan Lu, and Robert C. Williamson. 2026. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/corruptions-supervised-learning-problems-typology-and-mitigations\u0022 hreflang=\u0022en\u0022\u003ECorruptions of Supervised Learning Problems: Typology and Mitigations\u003C\/a\u003E\u201d. \u003Ci\u003EJournal of Machine Learning Research\u003C\/i\u003E 27 (72): 1-73.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/www.jmlr.org\/papers\/v27\/24-0808.html\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003ECorruption is notoriously widespread in data collection. Despite extensive research, the existing literature on corruption predominantly focuses on specific settings and learning scenarios, lacking a unified view. There is still a limited understanding of how to effectively model and mitigate corruption in machine learning problems. In this work, we develop a general theory of corruption from an information-theoretic perspective - with Markov kernels as a foundational mathematical tool. We generalize the definition of corruption beyond the concept of distributional shift: corruption includes all modifications of a learning problem, including changes in model class and loss function. We will focus here on changes in probability distributions. First, we construct a provably exhaustive framework for pairwise Markovian corruptions. The framework not only allows us to study corruption types based on their input space, but also serves to unify prior works on specific corruption models and establish a consistent nomenclature. Second, we systematically analyze the consequences of corruption on learning tasks by comparing Bayes risks in the clean and corrupted scenarios. This examination sheds light on complexities arising from joint and dependent corruptions on both labels and attributes. Notably, while label corruptions affect only the loss function, more intricate cases involving attribute corruptions extend the influence beyond the loss to affect the hypothesis class. Third, building upon these results, we investigate mitigations for various corruption types. We expand the existing loss-correction results for label corruption, and identify the necessity to generalize the classical corruption-corrected learning framework to a new paradigm with weaker requirements. Within the latter setting, we provide a negative result for loss correction in the attribute and the joint corruption case.\u003C\/p\u003E\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EH\u00f6ltgen, Benedikt, and Robert C. Williamson. 2026. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/costs-pretending-there-are-data-generating-probability-distributions-social-world\u0022 hreflang=\u0022en\u0022\u003EThe Costs of Pretending That There Are Data-Generating Probability Distributions in the Social World\u003C\/a\u003E\u201d. In \u003Ci\u003EFAccT \u201926: The 2026 ACM Conference on Fairness, Accountability, and Transparency\u003C\/i\u003E, 5219-36.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/dl.acm.org\/doi\/10.1145\/3805689.3806478\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EMachine Learning research, including work promoting fair or equitable algorithms, often relies on the concept of a data-generating probability distribution. The standard presumption is that since data points are \u0027sampled from\u0027 such a distribution, one can learn from observed data about this distribution and, thus, predict future data points which are also drawn from it. We argue, however, that such true probability distributions do not exist and that the rhetoric around them is harmful in social settings. We show that alternative frameworks focusing directly on relevant populations rather than abstract distributions are available and leave classical learning theory almost unchanged. Furthermore, we argue that the assumption of true probabilities or data-generating distributions can be misleading and obscure both the choices made and the goals pursued in machine learning practice. Based on these considerations, we suggest avoiding the assumption of data-generating probability distributions in the social world.\u003C\/p\u003E\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EWilliamson, Robert C. 2026. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/rhetoric\u0022 hreflang=\u0022en\u0022\u003EThe Rhetoric of Machine Learning\u003C\/a\u003E\u201d. \u003Ci\u003EArXiv 2604.06754 \u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022http:\/\/arxiv.org\/abs\/2604.06754\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EI examine the technology of machine learning from the perspective of rhetoric, which is simply the art of persuasion. Rather than being a neutral and \u0022objective\u0022 way to build \u0022world models\u0022 from data, machine learning is (I argue) inherently rhetorical. I explore some of its rhetorical features, and examine one pervasive business model where machine learning is widely used, \u0022manipulation as a service.\u0022\u003C\/p\u003E\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EWilliamson, Robert C. 2025. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/environmental-intelligence-context-and-rhetoric\u0022 hreflang=\u0022en\u0022\u003EEnvironmental Intelligence: Context and Rhetoric\u003C\/a\u003E\u201d. \u003Ci\u003EHarvard Data Science Review\u003C\/i\u003E 7 (4).\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/hdsr.mitpress.mit.edu\/pub\/5estwqmh\/download\/pdf\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n\n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EM\u00e9moli, Facundo, Brantley Vose, and Robert C. Williamson. (2025) 2025. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/geometry-and-stability-supervised-learning-problems\u0022 hreflang=\u0022en\u0022\u003EGeometry and Stability of Supervised Learning Problems\u003C\/a\u003E\u201d. \u003Ci\u003EJournal of Machine Learning Research\u003C\/i\u003E 26 (221): 1-99.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/jmlr.org\/beta\/papers\/v26\/24-0322.html\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp class=\u0022abstract mathjax\u0022\u003EWe introduce a notion of distance between supervised learning problems, which we call the Risk distance. This optimal-transport-inspired distance facilitates stability results; one can quantify how seriously issues like sampling bias, noise, limited data, and approximations might change a given problem by bounding how much these modifications can move the problem under the Risk distance. With the distance established, we explore the geometry of the resulting space of supervised learning problems, providing explicit geodesics and proving that the set of classification problems is dense in a larger class of problems. We also provide two variants of the Risk distance: one that incorporates specified weights on a problem\u0027s predictors, and one that is more sensitive to the contours of a problem\u0027s risk landscape.\u003C\/p\u003E\n\n\u003Cdiv class=\u0022metatable\u0022\u003E\n\u003Ctable\u003E\n\u003Ctbody\u003E\n\u003Ctr\u003E\n\t\u003Ctd class=\u0022tablecell label\u0022\u003E\u00a0\u003C\/td\u003E\n\u003C\/tr\u003E\n\u003C\/tbody\u003E\n\u003C\/table\u003E\n\u003C\/div\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EJohnston, David O., Cheng Soon Ong, and Robert C. Williamson. 2025. \u201c\u003Ca href=\u0022\/fmls\/publications\/decision-making-symmetry-and-structure-justifying-causal-interventions\u0022 hreflang=\u0022en\u0022\u003EDecision Making, Symmetry and Structure: Justifying Causal Interventions\u003C\/a\u003E\u201d. \u003Ci\u003EJournal of Causal Inference\u003C\/i\u003E 13.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/www.degruyter.com\/document\/doi\/10.1515\/jci-2023-0001\/html\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EWe can use structural causal models (SCMs) to help us evaluate the consequences of actions given data. SCMs identify actions with structural interventions. A careful decision maker may wonder whether this identification is justified. We seek such a justification. We begin with decision models, which map actions to distributions over outcomes but avoid additional causal assumptions. We then examine assumptions that could justify causal interventions, with a focus on symmetry. First, we introduce conditionally independent and identical responses (CIIR), a generalisation of the IID assumption to decision models. CIIR justifies identifying actions with interventions, but is often an implausible assumption. We consider an alternative: precedent is the assumption that\u201cwhat I can do has been done before, and its consequences observed,\u201d and is generally more plausible than CIIR. We show that precedent together with independence of causal mechanisms (ICM) and an observed conditional independence can justify identifying actions with causal interventions. ICM has been proposed as an alternative foundation for causal modelling, but this work suggests that it may in fact justify the interventional interpretation of causal models.\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EDerr, Rabanus, and Robert C Williamson. (2024) 2024. \u201c\u003Ca href=\u0022\/fmls\/publications\/fairness-and-randomness-machine-learning-statistical-independence-and-relativization\u0022 hreflang=\u0022en\u0022\u003EFairness and Randomness in Machine Learning: Statistical Independence and Relativization\u003C\/a\u003E\u201d. \u003Ci\u003EThe New England Journal of Statistics in Data Science \u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/doi.org\/10.51387\/24-NEJSDS73\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EFair Machine Learning endeavors to prevent unfairness arising in the context of machine learning applications embedded in society. Despite the variety of definitions of fairness and proposed \u0022fair algorithms\u0022, there remain unresolved conceptual problems regarding fairness. In this paper, we dissect the role of statistical independence in fairness and randomness notions regularly used in machine learning. Thereby, we are led to a suprising hypothesis: randomness and fairness can be considered equivalent concepts in machine learning.\u003Cbr\u003E\nIn particular, we obtain a relativized notion of randomness expressed as statistical independence by appealing to Von Mises\u0027 century-old foundations for probability. This notion turns out to be \u0022orthogonal\u0022 in an abstract sense to the commonly used i.i.d.-randomness. Using standard fairness notions in machine learning, which are defined via statistical independence, we then link the ex ante randomness assumptions about the data to the ex post requirements for fair predictions. This connection proves fruitful: we use it to argue that randomness and fairness are essentially relative and that both concepts should reflect their nature as modeling assumptions in machine learning.\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EWilliamson, Robert C. 2024. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/rhetoric-machine-learning\u0022 hreflang=\u0022en\u0022\u003EThe Rhetoric of Machine Learning\u003C\/a\u003E\u201d. In \u003Ci\u003E Persuasive Algorithms? A Symposium on the Rhetoric of Generative AI\u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003ETalk presented at the Persuasive Algorithms conference (which has no published proceedings)\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EFr\u00f6hlich, Christian, and Robert C. Williamson. 2024. \u201c\u003Ca href=\u0022\/fmls\/publications\/scoring-rules-and-calibration-imprecise-probabilities\u0022 hreflang=\u0022en\u0022\u003EScoring Rules and Calibration for Imprecise Probabilities\u003C\/a\u003E\u201d. \u003Ci\u003EArXiv\u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/arxiv.org\/abs\/2410.23001\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003E\u003Cspan class=\u0022abstract-full has-text-grey-dark mathjax\u0022 id=\u00222410.23001v1-abstract-full\u0022\u003EWhat does it mean to say that, for example, the probability for rain tomorrow is between 20% and 30%? The theory for the evaluation of precise probabilistic forecasts is well-developed and is grounded in the key concepts of proper scoring rules and calibration. For the case of imprecise probabilistic forecasts (sets of probabilities), such theory is still lacking. In this work, we therefore generalize proper scoring rules and calibration to the imprecise case. We develop these concepts as relative to data models and decision problems. As a consequence, the imprecision is embedded in a clear context. We establish a close link to the paradigm of (group) distributional robustness and in doing so provide new insights for it. We argue that proper scoring rules and calibration serve two distinct goals, which are aligned in the precise case, but intriguingly are not necessarily aligned in the imprecise case. The concept of decision-theoretic entropy plays a key role for both goals. Finally, we demonstrate the theoretical insights in machine learning practice, in particular we illustrate subtle pitfalls relating to the choice of loss function in distributional robustness. \u003C\/span\u003E\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EH\u00f6ltgen, Benedikt, and Robert C. Williamson. 2024. \u201c\u003Ca href=\u0022\/fmls\/publications\/which-distribution-were-you-sampled-towards-more-tangible-conception-data\u0022 hreflang=\u0022en\u0022\u003EWhich Distribution Were You Sampled From? Towards a More Tangible Conception of Data\u003C\/a\u003E\u201d. \u003Ci\u003EArXiv\u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/arxiv.org\/abs\/2407.17395\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003E\u003Cspan class=\u0022abstract-full has-text-grey-dark mathjax\u0022 id=\u00222407.17395v3-abstract-full\u0022\u003EMachine Learning research, as most of Statistics, heavily relies on the concept of a data-generating probability distribution. The standard presumption is that since data points are `sampled from\u0027 such a distribution, one can learn from observed data about this distribution and, thus, predict future data points which, it is presumed, are also drawn from it. Drawing on scholarship across disciplines, we here argue that this framework is not always a good model. Not only do such true probability distributions not exist; the framework can also be misleading and obscure both the choices made and the goals pursued in machine learning practice. We suggest an alternative framework that focuses on finite populations rather than abstract distributions; while classical learning theory can be left almost unchanged, it opens new opportunities, especially to model sampling. We compile these considerations into five reasons for modelling machine learning -- in some settings -- with finite populations rather than generative distributions, both to be more faithful to practice and to provide novel theoretical insights.\u003C\/span\u003E\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EH\u00f6ltgen, Benedikt, and Robert C. Williamson. 2024. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/causal-modelling-without-introducing-counterfactuals-or-abstract-distributions\u0022 hreflang=\u0022en\u0022\u003ECausal Modelling Without Introducing Counterfactuals or Abstract Distributions\u003C\/a\u003E\u201d. \u003Ci\u003EArXiv\u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/arxiv.org\/abs\/2407.17385\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003E\u003Cspan class=\u0022abstract-full has-text-grey-dark mathjax\u0022 id=\u00222407.17385v2-abstract-full\u0022\u003EThe most common approach to causal modelling is the potential outcomes framework due to Neyman and Rubin. In this framework, outcomes of counterfactual treatments are assumed to be well-defined. This metaphysical assumption is often thought to be problematic yet indispensable. The conventional approach relies not only on counterfactuals but also on abstract notions of distributions and assumptions of independence that are not directly testable. In this paper, we construe causal inference as treatment-wise predictions for finite populations where all assumptions are testable; this means that one can not only test predictions themselves (without any fundamental problem) but also investigate sources of error when they fail. The new framework highlights the model-dependence of causal claims as well as the difference between statistical and scientific inference.\u003C\/span\u003E\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003ERemeli, Mina, Moritz Hardt, and Robert C. Williamson. 2024. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/limits-predicting-online-speech-using-large-language-models\u0022 hreflang=\u0022en\u0022\u003ELimits to Predicting Online Speech Using Large Language Models\u003C\/a\u003E\u201d. \u003Ci\u003EArXiv\u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/arxiv.org\/abs\/2407.12850\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003E\u003Cspan class=\u0022abstract-full has-text-grey-dark mathjax\u0022 id=\u00222407.12850v1-abstract-full\u0022\u003EWe study the predictability of online speech on social media, and whether predictability improves with information outside a user\u0027s own posts. Recent work suggests that the predictive information contained in posts written by a user\u0027s peers can surpass that of the user\u0027s own posts. Motivated by the success of large language models, we empirically test this hypothesis. We define unpredictability as a measure of the model\u0027s uncertainty, i.e., its negative log-likelihood on future tokens given context. As the basis of our study, we collect a corpus of 6.25M posts from more than five thousand X (previously Twitter) users and their peers. Across three large language models ranging in size from 1 billion to 70 billion parameters, we find that predicting a user\u0027s posts from their peers\u0027 posts performs poorly. Moreover, the value of the user\u0027s own posts for prediction is consistently higher than that of their peers\u0027. Across the board, we find that the predictability of social media posts remains low, comparable to predicting financial news without context. We extend our investigation with a detailed analysis about the causes of unpredictability and the robustness of our findings. Specifically, we observe that a significant amount of predictive uncertainty comes from hashtags and @-mentions. Moreover, our results replicate if instead of prompting the model with additional context, we finetune on additional context.\u003C\/span\u003E\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EPacheco, Armando J. Cabrera, Rabanus Derr, and Robert C. Williamson. 2024. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/axiomatic-approach-loss-aggregation-and-adapted-aggregating-algorithm\u0022 hreflang=\u0022en\u0022\u003EAn Axiomatic Approach to Loss Aggregation and an Adapted Aggregating Algorithm\u003C\/a\u003E\u201d. \u003Ci\u003EArXiv\u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/arxiv.org\/abs\/2406.02292\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003E\u003Cspan class=\u0022abstract-full has-text-grey-dark mathjax\u0022 id=\u00222406.02292v1-abstract-full\u0022\u003ESupervised learning has gone beyond the expected risk minimization framework. Central to most of these developments is the introduction of more general aggregation functions for losses incurred by the learner. In this paper, we turn towards online learning under expert advice. Via easily justified assumptions we characterize a set of reasonable loss aggregation functions as quasi-sums. Based upon this insight, we suggest a variant of the Aggregating Algorithm tailored to these more general aggregation functions. This variant inherits most of the nice theoretical properties of the AA, such as recovery of Bayes\u0027 updating and a time-independent bound on quasi-sum regret. Finally, we argue that generalized aggregations express the attitude of the learner towards losses.\u003C\/span\u003E\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EFr\u00f6hlich, Christian, Rabanus Derr, and Robert C Williamson. 2024. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/strictly-frequentist-imprecise-probability\u0022 hreflang=\u0022en\u0022\u003EStrictly Frequentist Imprecise Probability\u003C\/a\u003E\u201d. \u003Ci\u003EInternational Journal of Approximate Reasoning\u003C\/i\u003E 168 (109148).\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/www.sciencedirect.com\/science\/article\/pii\/S0888613X24000355\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EStrict frequentism defines probability as the limiting relative frequency in an infinite se- quence. What if the limit does not exist? We present a broader theory, which is applicable also to random phenomena that exhibit diverging relative frequencies. In doing so, we develop a close connection with the theory of imprecise probability: the cluster points of relative frequencies yield an upper probability. We show that a natural frequentist definition of conditional probability recovers the generalized Bayes rule. This also sug- gests an independence concept, which is related to epistemic irrelevance in the imprecise probability literature. Finally, we prove constructively that, for a finite set of elementary events, there exists a sequence for which the cluster points of relative frequencies coincide with a prespecified set which demonstrates the naturalness, and arguably completeness, of our theory.\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EFr\u00f6hlich, Christian, and Robert C Williamson. (2024) 2024. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/risk-measures-and-upper-probabilities-coherence-and-stratification\u0022 hreflang=\u0022en\u0022\u003ERisk Measures and Upper Probabilities: Coherence and Stratification\u003C\/a\u003E\u201d. \u003Ci\u003EJournal of Machine Learning Research\u003C\/i\u003E 25 (207): 1-100.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022http:\/\/jmlr.org\/papers\/v25\/22-0641.html\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EMachine learning typically presupposes classical probability theory which implies that aggregation is built upon expectation. There are now multiple reasons to motivate looking at richer alternatives to classical probability theory as a mathematical foundation for machine learning. We systematically examine a powerful and rich class of such alternatives, known variously as spectral risk measures, Choquet integrals or Lorentz norms. We present a range of characterization results, and demonstrate what makes this spectral family so special. In doing so we demonstrate a natural stratification of all coherent risk measures in terms of the upper probabilities that they induce by exploiting results from the theory of rearrangement invariant Banach spaces. We empirically demonstrate how this new approach to uncertainty helps tackling practical machine learning problems.\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EFr\u00f6lich, Christian, and Robert C. Williamson. 2024. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/data-models-two-manifestations-imprecision\u0022 hreflang=\u0022en\u0022\u003EData Models With Two Manifestations of Imprecision\u003C\/a\u003E\u201d. \u003Ci\u003EArXiv\u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/arxiv.org\/abs\/2404.09741\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EMotivated by recently emerging problems in machine learning and statistics, we propose data models which relax the familiar i.i.d. assumption. In essence, we seek to understand what it means for data to come from a set of probability measures. We show that our frequentist data models, parameterized by such sets, manifest two aspects of imprecision. We characterize the intricate interplay of these manifestations, aggregate (ir)regularity and local (ir)regularity, where a much richer set of behaviours compared to an i.i.d. model is possible. In doing so we shed new light on the relationship between non-stationary, locally precise and stationary, locally imprecise data models. We discuss possible applications of these data models in machine learning and how the set of probabilities can be estimated. For the estimation of aggregate irregularity, we provide a negative result but argue that it does not warrant pessimism. Understanding these frequentist aspects of imprecise probabilities paves the way for deriving generalization of proper scoring rules and calibration to the imprecise case, which can then contribute to tackling practical problems.\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EWilliamson, Robert C, and Zac Cranko. 2024. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/information-processing-equalities-and-information-risk-bridge\u0022 hreflang=\u0022en\u0022\u003EInformation Processing Equalities and the Information-Risk Bridge\u003C\/a\u003E\u201d. \u003Ci\u003EJournal of Machine Learning Research\u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022http:\/\/jmlr.org\/papers\/v25\/22-0988.html\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cdiv class=\u0022tex2jax_process\u0022\u003E\u003Cp\u003EWe introduce two new classes of measures of information for statistical experiments which generalise and subsume \u003Cspan class=\u0022math-tex\u0022\u003E\\(\\phi\\)\u003C\/span\u003E-divergences, integral probability metrics, \u003Cspan class=\u0022math-tex\u0022\u003E\\(\\mathfrak{N}\\)\u003C\/span\u003E-distances (MMD), and \u003Cspan class=\u0022math-tex\u0022\u003E\\((f,\\Gamma)\\)\u003C\/span\u003E-divergences between two or more distributions. This enables us to derive a simple geometrical relationship between measures of information and the Bayes risk of a statistical decision problem, thus extending the variational \u003Cspan class=\u0022math-tex\u0022\u003E\\(\\phi\\)\u003C\/span\u003E-divergence representation to multiple distributions in an entirely symmetric manner. The new families of divergence are closed under the action of Markov operators which yields an information processing equality which is a refinement and generalisation of the classical data processing inequality. This equality gives insight into the significance of the choice of the hypothesis class in classical risk minimization.\u003C\/p\u003E\n\n\u003Cp\u003E\u003Cspan style=\u0022color:#c0392b;\u0022\u003E\u003Cem\u003EAll Information\u003Cbr\u003E\nUndergoes transformation.\u003Cbr\u003E\nNo gap. Now equal.\u003C\/em\u003E\u003C\/span\u003E\u003C\/p\u003E\n\u003C\/div\u003E\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EDerr, Rabanus, and Robert C. Williamson. 2024. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/four-facets-forecast-felicity-calibration-predictiveness-randomness-and-regret\u0022 hreflang=\u0022en\u0022\u003EFour Facets of Forecast Felicity: Calibration, Predictiveness, Randomness and Regret\u003C\/a\u003E.\u201d\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/arxiv.org\/pdf\/2401.14483.pdf\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EMachine learning is about forecasting. Forecasts, however, obtain their usefulness only through their evaluation. Machine learning has traditionally focused on types of losses and their corresponding regret. Currently, the machine learning community regained interest in calibration. In this work, we show the conceptual equivalence of calibration and regret in evaluating forecasts. We frame the evaluation problem as a game between a forecaster, a gambler and nature. Putting intuitive restrictions on gambler and forecaster, calibration and regret naturally fall out of the framework. In addition, this game links evaluation of forecasts to randomness of outcomes. Random outcomes with respect to forecasts are equivalent to good forecasts with respect to outcomes. We call those dual aspects, calibration and regret, predictiveness and randomness, the four facets of forecast felicity.\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EWilliamson, Robert C, and Zac Cranko. (2023) 2023. \u201c\u003Ca href=\u0022\/fmls\/publications\/geometry-and-calculus-losses\u0022 hreflang=\u0022en\u0022\u003EThe Geometry and Calculus of Losses\u003C\/a\u003E\u201d. \u003Ci\u003EJournal of Machine Learning Research\u003C\/i\u003E 24 (342): 1-72.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/www.jmlr.org\/papers\/v24\/22-0987.html\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EStatistical decision problems are the foundation of statistical machine learning. The simplest problems are binary and multiclass classification and class probability estimation. Central to their definition is the choice of loss function, which is the means by which the quality of a solution is evaluated. In this paper we systematically develop the theory of loss functions for such problems from a novel perspective whose basic ingredients are convex sets with a particular structure. The loss function is defined as the subgradient of the support function of the convex set. It is consequently automatically proper (calibrated for probability estimation). This perspective provides three novel opportunities. It enables the development of a fundamental relationship between losses and (anti)-norms that appears to have not been noticed before. Second, it enables the development of a calculus of losses induced by the calculus of convex sets which allows the interpolation between different losses, and thus is a potential useful design tool for tailoring losses to particular problems. In doing this we build upon, and considerably extend, existing results on M-sums of convex sets. Third, the perspective leads to a natural theory of `polar\u0027 (or `inverse\u0027) loss functions, which are derived from the polar dual of the convex set defining the loss, and which form a natural universal substitution function for Vovk\u0027s aggregating algorithm.\u003C\/p\u003E\n\n\u003Cp\u003E\u003Cspan style=\u0022color:#c0392b;\u0022\u003E\u003Cem\u003ESupport gradients,\u003Cbr\u003E\ncontrol proper loss functions:\u003Cbr\u003E\nSecretly convex.\u003C\/em\u003E\u003C\/span\u003E\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003ECabrera.Pacheco, Armando J., and Robert C Williamson. 2023. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/geometry-mixability\u0022 hreflang=\u0022en\u0022\u003EThe Geometry of Mixability\u003C\/a\u003E\u201d. \u003Ci\u003ETransactions on Machine Learning Research\u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/openreview.net\/forum?id=VrvGHDSzZ7\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EMixable loss functions are of fundamental importance in the context of prediction with expert advice in the online setting since they characterize fast learning rates. By re-interpreting properness from the point of view of differential geometry, we provide a simple geometric characterization of mixability for the binary and multi-class cases: a proper loss function l is \u03b7-mixable if and only if the superpredition set spr(\u03b7l) of the scaled loss function \u03b7l slides freely inside the superprediction set spr(llog) of the log loss llog, under fairly general assumptions on the differentiability of l. Our approach provides a way to treat some concepts concerning loss functions (like properness) in a \u201ccoordinate-free\u201d manner and reconciles previous results obtained for mixable loss functions for the binary and the multi-class cases.\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EDerr, Rabanus, and Robert C. Williamson. (2023) 2023. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/systems-precision-coherent-probabilities-pre-dynkin-systems-and-coherent-previsions\u0022 hreflang=\u0022en\u0022\u003ESystems of Precision: Coherent Probabilities on Pre-Dynkin Systems and Coherent Previsions on Linear Subspaces\u003C\/a\u003E\u201d. \u003Ci\u003EEntropy\u003C\/i\u003E 25 (9): 1283. https:\/\/doi.org\/10.3390\/e25091283 .\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/www.mdpi.com\/1099-4300\/25\/9\/1283\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EAbstract\u003C\/p\u003E\n\n\n\u003Cdiv class=\u0022html-article-content\u0022\u003E\n\u003Cdiv class=\u0022html-dynamic\u0022\u003E\n\n\u003Cdiv class=\u0022art-abstract art-abstract-new in-tab hypothesis_container\u0022\u003E\n\u003Cdiv\u003E\n\n\u003Cdiv class=\u0022html-p\u0022\u003EIn the literature on imprecise probability, little attention is paid to the fact that imprecise probabilities are precise on a set of events. We call these sets \u003Cem\u003E\u003Cspan class=\u0022html-italic\u0022\u003Esystems of precision\u003C\/span\u003E\u003C\/em\u003E. We show that, under mild assumptions, the system of precision of a lower and upper probability form a so-called (pre-)Dynkin system. Interestingly, there are several settings, ranging from machine learning on partial data over frequential probability theory to quantum probability theory and decision making under uncertainty, in which, a priori, the probabilities are only desired to be precise on a specific underlying set system. Here, (pre-)Dynkin systems have been adopted as systems of precision, too. We show that, under extendability conditions, those pre-Dynkin systems equipped with probabilities can be embedded into algebras of sets. Surprisingly, the extendability conditions elaborated in a strand of work in quantum probability are equivalent to coherence from the imprecise probability literature. On this basis, we spell out a lattice duality which relates systems of precision to credal sets of probabilities. We conclude the presentation with a generalization of the framework to expectation-type counterparts of imprecise probabilities. The analogue of pre-Dynkin systems turns out to be (sets of) linear subspaces in the space of bounded, real-valued functions. We introduce partial expectations, natural generalizations of probabilities defined on pre-Dynkin systems. Again, coherence and extendability are equivalent. A related but more general lattice duality preserves the relation between systems of precision and credal sets of probabilities.\u003C\/div\u003E\n\n\u003C\/div\u003E\n\u003C\/div\u003E\n\n\u003C\/div\u003E\n\u003C\/div\u003E\n\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EMansour, Yishay, Richard Nock, and Robert C. Williamson. 2023. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/random-classification-noise-does-not-defeat-all-convex-potential-boosters-irrespective\u0022 hreflang=\u0022en\u0022\u003ERandom Classification Noise Does Not Defeat All Convex Potential Boosters Irrespective of Model Choice\u003C\/a\u003E\u201d. In \u003Ci\u003EICML2023 \u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/openreview.net\/forum?id=1UaGAhLAsL\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EA landmark negative result of Long and Servedio has had a considerable impact on research and development in boosting algorithms, around the now famous tagline that \u0022noise defeats all convex boosters\u0022. In this paper, we appeal to the half-century+ founding theory of losses for class probability estimation, an extension of Long and Servedio\u0027s results and a new general convex booster to demonstrate that the source of their negative result is in fact the \u003Cem\u003Emodel class\u003C\/em\u003E, linear separators. Losses or algorithms are neither to blame. This leads us to a discussion on an otherwise praised aspect of ML, \u003Cem\u003Eparameterisation\u003C\/em\u003E.\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EFr\u00f6hlich, Christian, and Robert C. Williamson. 2023. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/insights-insurance-fair-machine-learning-responsibility-performativity-and-aggregates\u0022 hreflang=\u0022en\u0022\u003EInsights From Insurance for Fair Machine Learning: Responsibility, Performativity and Aggregates\u003C\/a\u003E\u201d. \u003Ci\u003EArXiv\u003C\/i\u003E. https:\/\/doi.org\/10.48550\/arXiv.2306.14624.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/arxiv.org\/abs\/2306.14624\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EWe argue that insurance can act as an analogon for the social situatedness of machine learning systems, hence allowing machine learning scholars to take insights from the rich and interdisciplinary insurance literature. Tracing the interaction of uncertainty, fairness and responsibility in insurance provides a fresh perspective on fairness in machine learning. We link insurance fairness conceptions to their machine learning relatives, and use this bridge to problematize fairness as calibration. In this process, we bring to the forefront three themes that have been largely overlooked in the machine learning literature: responsibility, performativity and tensions between aggregate and individual.\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EH\u00f6ltgen, Benedikt, and Robert C Williamson. 2023. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/richness-calibration\u0022 hreflang=\u0022en\u0022\u003EOn the Richness of Calibration\u003C\/a\u003E\u201d. In \u003Ci\u003EFAccT2023 \u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/orange.hosting.lsoft.com\/trk\/clickp?ref=znwrbbrs9_6-2d8c7_0x33ae25x0421\u0026amp;doi=3593013.3594068\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EProbabilistic predictions can be evaluated through comparisons with observed label frequencies, that is, through the lens of calibration. Recent scholarship on algorithmic fairness has started to look at a growing variety of calibration-based objectives under the name of multi-calibration but has still remained fairly restricted. In this paper, we explore and analyse forms of evaluation through calibration by making explicit the choices involved in design- ing calibration scores. We organise these into three grouping choices and a choice concerning the agglomeration of group errors. This provides a framework for comparing previously proposed calibration scores and helps to formulate novel ones with desirable mathematical properties. In particular, we explore the possibility of grouping datapoints based on their input features rather than on predictions and formally demonstrate advantages of such approaches. We also characterise the space of suitable agglomeration functions for group errors, generalising previously proposed calibration scores. Complementary to such population-level scores, we explore calibration scores at the individual level and analyse their relationship to choices of grouping. We draw on these insights to introduce and axiomatise fairness deviation measures for population-level scores. We demonstrate that with appropriate choices of grouping, these novel global fairness scores can provide notions of (sub-)group or individual fairness.\u003C\/p\u003E\n\n\u003Cp\u003E\u003Cem\u003E\u003Cspan style=\u0022color:#e74c3c;\u0022\u003E\u003Cspan style=\u0022font-family:\u0027AvenirNext-Regular\u0027;font-size:14px;font-style:normal;font-weight:400;letter-spacing:normal;text-indent:0px;text-transform:none;white-space:normal;word-spacing:0px;text-decoration:none;float:none;\u0022\u003ESo many options,\u003C\/span\u003E\u003Cbr style=\u0022color:rgb(0,0,0);font-family:\u0027AvenirNext-Regular\u0027;font-size:14px;font-style:normal;font-weight:400;letter-spacing:normal;text-indent:0px;text-transform:none;white-space:normal;word-spacing:0px;text-decoration:none;\u0022\u003E\n\u003Cspan style=\u0022font-family:\u0027AvenirNext-Regular\u0027;font-size:14px;font-style:normal;font-weight:400;letter-spacing:normal;text-indent:0px;text-transform:none;white-space:normal;word-spacing:0px;text-decoration:none;float:none;\u0022\u003Eso much potential, hiding\u003C\/span\u003E\u003Cbr style=\u0022color:rgb(0,0,0);font-family:\u0027AvenirNext-Regular\u0027;font-size:14px;font-style:normal;font-weight:400;letter-spacing:normal;text-indent:0px;text-transform:none;white-space:normal;word-spacing:0px;text-decoration:none;\u0022\u003E\n\u003Cspan style=\u0022font-family:\u0027AvenirNext-Regular\u0027;font-size:14px;font-style:normal;font-weight:400;letter-spacing:normal;text-indent:0px;text-transform:none;white-space:normal;word-spacing:0px;text-decoration:none;float:none;\u0022\u003Ein calibration.\u003C\/span\u003E\u003C\/span\u003E\u003C\/em\u003E\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EDerr, Rabanus, and Robert C Williamson. 2023. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/set-structure-precision-coherent-probabilities-pre-dynkin-systems\u0022 hreflang=\u0022en\u0022\u003EThe Set Structure of Precision: Coherent Probabilities on Pre-Dynkin-Systems\u003C\/a\u003E\u201d. \u003Ci\u003EArXiv\u003C\/i\u003E arXiv:2302.03522.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003E\u003Cbr\u003E\nIn literature on imprecise probability little attention is paid to the fact that imprecise proba- bilities are precise on some events. We call these sets system of precision. We show that, under mild assumptions, the system of precision of a lower and upper probability form a so-called (pre-)Dynkin-system. Interestingly, there are several settings, ranging from machine learning on partial data over frequential probability theory to quantum probability theory and decision making under uncertainty, in which a priori the probabilities are only desired to be precise on a specific underlying set system. At the core of all of these settings lies the observation that precise beliefs, probabilities or frequencies on two events do not necessarily imply this precision to hold for the intersection of those events. Here, (pre-)Dynkin-systems have been adopted as systems of precision, too. We show that, under extendability conditions, those pre-Dynkin- systems equipped with probabilities can be embedded into algebras of sets. Surprisingly, the extendability conditions elaborated in a strand of work in quantum physics are equivalent to coherence in the sense of Walley [Walley, 1991, p. 84]. Thus, literature on probabilities on pre-Dynkin-systems gets linked to the literature on imprecise probability. Finally, we spell out a lattice duality which rigorously relates the system of precision to credal sets of probabilities. In particular, we provide a hitherto undescribed, parametrized family of coherent imprecise probabilities.\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EFr\u00f6hlich, Christian, and Robert C Williamson. 2023. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/tailoring-tails-risk-measures-fine-grained-tail-sensitivity\u0022 hreflang=\u0022en\u0022\u003ETailoring to the Tails: Risk Measures for Fine-Grained Tail Sensitivity\u003C\/a\u003E\u201d. \u003Ci\u003ETransactions on Machine Learning Research\u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n\n  \u003Cdiv class=\u0022field field--name-publishers-version field--type-link field--label-visually_hidden field--mode-teaser\u0022\u003E\n    \u003Cdiv class=\u0022field--label sr-only\u0022\u003EPublisher\u0027s Version\u003C\/div\u003E\n              \u003Cdiv class=\u0022field--item\u0022\u003E\u003Ca href=\u0022https:\/\/openreview.net\/pdf?id=UntUoeLwwu\u0022\u003EPublisher\u0026#039;s Version\u003C\/a\u003E\u003C\/div\u003E\n          \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EExpected risk minimization (ERM) is at the core of machine learning systems. This means that the risk inherent in a loss distribution is summarized using a single number - its average. In this paper, we propose a general approach to construct risk measures which exhibit a desired tail sensitivity and may replace the expectation operator in ERM. Our method relies on the specification of a reference distribution with a desired tail behaviour, which is in a one-to-one correspondence to a coherent upper probability. Any risk measure, which is compatible with this upper probability, displays a tail sensitivity which is finely tuned to the reference distribution. As a concrete example, we focus on divergence risk measures based on f-divergence ambiguity sets, which are a widespread tool used to foster distributional robustness of machine learning systems. For instance, we show how ambiguity sets based on the Kullback-Leibler divergence are intricately tied to the class of subexponential random variables. We elaborate the connection of divergence risk measures and rearrangement invariant Banach norms.\u003C\/p\u003E\n\n\u003Cp\u003E\u003Cbr\u003E\n\u003Cspan style=\u0022color:#c0392b;\u0022\u003E\u003Cem\u003ERare loss. Control how?\u003Cbr\u003E\nThe fundamental function.\u003Cbr\u003E\nYour tails are tailored.\u003C\/em\u003E\u003C\/span\u003E\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n \u003Cli\u003E\u003Carticle class=\u0022bibcite-reference\u0022\u003E\n  \n  \n      \u003Cdiv class=\u0022bibcite-citation\u0022\u003E\n      \u003Cdiv class=\u0022csl-bib-body\u0022\u003E\u003Cdiv class=\u0022csl-entry\u0022\u003EMansour, Yishay, Richard Nock, and Robert C Williamson. 2022. \u201c\u003Ca href=\u0022https:\/\/fm.ls\/publications\/what-killed-convex-booster\u0022 hreflang=\u0022en\u0022\u003EWhat Killed the Convex Booster?\u003C\/a\u003E\u201d. \u003Ci\u003EArXiv Preprint ArXiv:2205.09628\u003C\/i\u003E.\u003C\/div\u003E\u003C\/div\u003E\n  \u003C\/div\u003E\n                \u003Cdiv class=\u0022field--label field--abstract\u0022\u003E\n      \u003Cbutton class=\u0022btn-abstract collapsed\u0022 data-toggle=\u0022collapse\u0022 data-target=\u0022#collapseAbstract\u0022 aria-expanded=\u0022false\u0022 aria-controls=\u0022collapseAbstract\u0022\u003EAbstract \u003C\/button\u003E\n    \u003C\/div\u003E\n                  \u003Cdiv class=\u0022field--item abstract--content collapse\u0022 id=\u0022collapseAbstract\u0022 aria-expanded=\u0026quot;false\u0026quot;\u003E\u003Cp\u003EA landmark negative result of Long and Servedio established a worst-case spectacular failure of a supervised learning trio (loss, algorithm, model) otherwise praised for its high precision machinery. Hundreds of papers followed up on the two suspected culprits: the loss (for being convex) and\/or the algorithm (for fitting a classical boosting blueprint). Here, we call to the half-century+ founding theory of losses for class probability estimation (properness), an extension of Long and Servedio\u0027s results and a new general boosting algorithm to demonstrate that the real culprit in their specific context was in fact the (linear) model class. We advocate for a more general standpoint on the problem as we argue that the source of the negative result lies in the dark side of a pervasive -- and otherwise prized -- aspect of ML: \u003Cem\u003Eparameterisation\u003C\/em\u003E.\u003C\/p\u003E\n\u003C\/div\u003E\n        \n  \u003C\/article\u003E\n\u003C\/li\u003E\n\u003C\/ul\u003E\n  \u003Cnav role=\u0022navigation\u0022 aria-labelledby=\u0022pagination-for-lop-publications\u0022 id=pager-heading\u003E\n    \u003Ch3 id=\u0022pagination-for-lop-publications\u0022 class=\u0022visually-hidden\u0022\u003Epagination for lop publications\u003C\/h3\u003E\n    \u003Cul class=\u0022js-pager__items pager-mini\u0022\u003E\n            \u003Cli class=\u0022current\u0022\u003E\n        \u003Cspan aria-live=\u0022polite\u0022\u003E\n            \u003Cspan class=\u0022visually-hidden\u0022\u003ELOP - Publications\u003C\/span\u003E\n            1 of 2\n          \u003C\/span\u003E      \u003C\/li\u003E\n              \u003Cli\u003E\n          \u003Ca href=\u0022https:\/\/fm.ls\/refresh-widget-content\/6158?page=1\u0026amp;selector=list-of-posts\u0026amp;pagerid=pager-heading\u0026amp;moreid=node-readmore\u0022 class=\u0022use-ajax next\u0022 rel=\u0022next\u0022\u003E\u003Cspan aria-hidden=\u0022true\u0022\u003E\u203a\u203a\u003C\/span\u003E\u003Cspan class=\u0022visually-hidden\u0022\u003ENext page\u003C\/span\u003E\u003C\/a\u003E\n        \u003C\/li\u003E\n          \u003C\/ul\u003E\n  \u003C\/nav\u003E\n\n\u003Cdiv class=\u0022node-readmore\u0022 id=node-readmore\u003E\u003C\/div\u003E\n","settings":null},{"command":"insert","method":"replaceWith","selector":"#","data":"","settings":null},{"command":"insert","method":"replaceWith","selector":"#","data":"","settings":null},{"command":"insert","method":"replaceWith","selector":".field--name-field-widget-title","data":"","settings":null}]