A personal update on infinite ethics
It's not as bad as I thought.
(Epistemic status: quickly written, opinionated)
I’m on the record as being pretty worried about the problem:
I was pretty disturbed by Joe Carlsmith’s grave warnings about it in 2022 - it made the issue look extremely thorny and difficult to resolve (if you don’t know this area at all, you should read his piece - it’s a breezy read and a great introduction).
But I just came across a paper by Toby Ord from last year that I somehow missed: Evaluating the infinite. I think people have tried things in this direction before - but he goes farther mathematically, and it’s the first time I’ve seen it laid out in such a convincing way, with all its advantages made clear.
He uses the hyperreal numbers. I’m not going to explain them here, you should just read the paper - it’s really nice.
The most important thing is that in the hyperreal context, infinity plus one does not equal infinity. So this resolves the deepest concern, “infinitarian paralysis” - even if the world is infinitely large, finite consequences still matter.
And it has many other nice properties, including (roughly speaking):
Strong Pareto: If A and B are identical except A is better than B on some elements, A is better than B.
Finite anonymity: Reordering a finite number of elements doesn’t change anything.
Carlsmith actually briefly addresses the hyperreal approach in his essay (in section 8), but quickly dismisses it - and his objections seem a little weak-sauce after reading the Ord paper.1
That said, my credence on completely different solutions is definitely not zero2. There are two big issues for me (that actually feel like conceptual confusions and not “just” implementation difficulties):
Infinite anonymity (“Reordering an infinite number of elements doesn’t change anything”) is not met. In fact, Strong Pareto is incompatible with it in general. I’m really confused about how much this matters in practice - do any of our actions result in an infinite reordering? I guess so, if we start an action one time-step later while holding everything else equal? But that’s not really possible, is it? And for the one example of infinite reordering changing something that Ord gives3, there actually is a reasonable-seeming intuition. Is that always the case?
Also, an argument that this is fine: a big point that Ord makes is that cardinal numbers (with their condition of “if there’s a bijection, it’s equal”) are too coarse-grained to distinguish different infinities. So it makes sense that Infinite Anonymity would not be met - rearranging a sequence is just putting it through a bijection with itself. It’s the same bijection intuition that needs to be stamped out - hyperreal infinities are nuanced enough that the “order/structure” of things matter, not just their “raw amount”.When summing over infinite space, the result depends on the point of origin for our integration (which again makes sense, since switching point of origin is basically an infinite reordering of space). That seems annoying - I guess you could just make it dependent on the observer, but that seems unprincipled. Again, I’m confused about how big of an issue this actually is in practice.
But it seems like a really really great start - I’m kind of convinced by the minimal takeaway of “stop using cardinals, it makes no sense, distinguish different kinds of infinities, e.g. with hyperreals”. It’s also an “existence proof” that a principled approach doesn’t necessarily immediately run into complete nonsense.
That feels very important, similar to when I first read the famous FDT paper. It moves me from “Oh god, I have no idea what’s going on” to “Okay, it seems like this rough approach seems plausible and I don’t need to radically change my worldview”.4
For example, these quotes from Carlsmith seem too strong and overconfident to me now:5
My point is just that [the] response isn’t going to look like the simple, complete, neutrality-respecting, totalist, hedonistic, EV-maximizing utilitarianism that some hoped, back in the day, would answer every ethical question
And to be clear: I don’t think that understanding the ethics, here, is going to look like “patching a few counterexamples to expansionism” or “figuring out how to deal with lotteries involving incomparable outcomes.” I’m imagining something closer to: “understanding ~all the math you might ever need, including everything related to all the infinites on the completed version of that crazy chart above; solving all of cosmology, physics, metaphysics, epistemology, and so on, too; probably reconceptualizing everything in fundamentally new and more sophisticated terms — terms that creatures at our current level of cognitive capacity can’t grok; then building up a comprehensive ethics and decision theory (assuming those terms still make sense), informed by this understanding, and encompassing of all the infinities that this understanding makes relevant.” It may well make sense to get started on this project now (or it might not); but we’re not, as it were, a few papers away.
I’m happy that my reaction back in 2022 was not to abandon “the utilitarian dream”, as Carlsmith calls it6 - I thought “Okay, this is pretty weird, but there has to be some kind of solution - the intuitions pushing for some kind of basic total consequentialism are just too powerful. I don’t want to bound my utility, that seems very unnatural. There’s probably some kind of new math that’s needed, some different way of thinking about infinities.”
I was right, I think!
EDIT: Here is a nice chat with Opus 4.8 about some of the math and its limitations that clarified some things for me.7
EDIT 2: I want to make something explicit that Ord left implicit but that IMO is a big deal, and arguably a main takeaway of the paper (going to be speaking kind of loosely here): All the main “downsides” of the approach are actually just necessary consequences of accepting that cardinality is not the right measure of size:
Infinite permutations are just bijections (i.e. how you compare sizes in the cardinal approach), so it makes sense that they wouldn’t generally preserve size here.
Choosing a different point of origin for integration is just an infinite permutation of space, so it makes sense that it wouldn’t generally preserve size.
Changing the order of integration of two variables in an improper integral is just an infinite permutation (and a drastic one at that), so it makes sense that nested/cartesian integrals don’t preserve size and that we need to do spherical/”expanding volume” integrals in multidimensional settings.8
This intuitively makes the whole approach even more convincing - we are not running into random difficulties, but just the same core insight over and over again.
(And to further summarize my current understanding (I believe this was true even before this paper, Ord just gives a concrete approach): I think the case for this philosophical point is also pretty strong - our main concern, infinitarian paralysis, is intrinsic to the cardinal approach. We literally can’t use cardinal infinities (i.e., in practice, the extended reals) and not have infinitarian paralysis - cardinals necessarily mean that infinity plus one equals infinity.
So if we don’t want infinitarian paralysis (and we want to take infinity seriously, i.e. not do discounting or bounding or averaging), it seems like we have to accept these downsides above, and this doesn’t even depend on the specific choice of hyperreals as the alternative.9
And the downsides don’t even seem that bad, much less bad than other counterintuitive things about reality I’ve accepted in quantum physics or general relativity or decision theory. I can see how there might be some kind of intuition that of course you can’t just magically garble an infinite sequence however you want, change the relative frequencies of everything, and preserve its sum - bijection is literally you doing whatever you want to the sequence, it’s not even close to structure-preserving enough to be our guiding light.
Hm, yeah, something like, very speculatively and vaguely… “In finite contexts, we can ignore the structure of a sequence because all the elements “collapse” together when the summing process terminates, and all information about the order/structure is destroyed. Only in the partial sums does the order/structure have an effect. But an infinite summing process never terminates, we are forever in a partial sum, and so the structure of the sequence is extremely important to describing the infinite sum. Therefore, we cannot rely on the bijection as our tool of comparison, since it is literally the most permissive, structurally lossy permutation possible”.
EDIT 3: Here’s something Ord said in an email exchange that I found useful (reproducing it with permission):
Part of the problem is that maths education has now instilled in many of us the counterintuitive principles of Hilbert’s hotel and the cardinal numbers, and that while these are a very useful concept of infinity, they are not the right one for this job and so our current mathematical intuitions are leading us more astray than if we were mathematically naive.
Yeah, again, my main takeaway is not even about hyperreals specifically, but that we should start acting like infinity + 1 > infinity. It feels like some kind of historical accident that it seems so obvious to us all that infinity + 1 = infinity. What were we thinking?
Getting more into the weeds here: Specifically, the arbitrariness of the ultra-filter doesn’t seem as bad after reading about Ord’s potential solution of averaging between them (although I don’t fully understand it), and the examples of ultra-filter dependence he gives actually seem intuitively unclear as well ( (0,0,0,…) vs (1,-2,1,1,-2,1,1,-2,…) ) (so maybe it’s good if our math reflects that) (credit for this point to Bostrom), contrived (“a sequence that takes every finite value infinitely many times”), or not that bad (a constant difference).
Especially that infinity is not a coherent concept, perhaps because it’s uncomputable or something. There was a great LW post about this years ago that neither me nor Opus 4.8 can find again.
(0,1,2,3,…) vs (1,2,3,…)
Well, acausal considerations do change things quite a lot - but not as far as needing to abandon total consequentialism. And I wouldn’t say it goes as far as the FDT paper deconfusion - maybe 2/3rds of the way there.
Concretely, I don’t think the presence of many weird issues should update you this much about the difficulty of the domain - there are plenty of examples in the history of science of paradigm shifts resolving many seemingly unrelated problems at once.
A non-naive version of it, of course. Even in a finite setting, we face judgment calls where there is probably no physical or theoretical fact of the matter (e.g. whether moments of unsettled ecstasy or calm bliss are more valuable). And there are acausal cooperative values and virtue ethics as a strong heuristic in the face of deep uncertainty, and so on.
I’m left with the rough impression that the “averaging over ultrafilters” thing generally works for “non-weird” sequences, and that averaging over different points of origin seems promising as well?
And this also links up nicely with expansionism, another major approach to infinite ethics.
I think they also apply to expansionism and the overtaking criterion, although I’m not sure.



The biggest problem imo is that it implies that you might improve things by making everyone worse off and moving people around.
I prefer a solution that holds that there’s ubiquitous incompatibility but still uses hypereeals to evaluate effects of actions
I think order dependence is a serious problem - where would you get an order from, thats not ultimately space-time? Because once you admit space-time dependence, its not clear what the motivation even is doing it with summation. Hyperreals make a lot of comparison worth taking a look at and wondering if you really do think thats better. For example, consider the utility streams:
⟨1,-1,2,-2,...⟩ vs ⟨-1,1,-2,2,...⟩
and
⟨0,1,-1,2,-2,...⟩ vs ⟨0,-1,1,-2,2,...⟩
One of these is 0 vs 0, and the other is +ω/2 vs -ω/2 (which is which depends on the ultrafilter). Basically everything about this is bad: Putting a zero in front of two options changes the comparison. We're drawing infinite value differences from a difference in ordering thats trivially created by different points of origen. And... the things thats better about the better option here is apparently just that the better things happen earlier. So, in order to avoid temporal discounting, which would have found a finite difference between these worlds, we have adopted a theory which finds an infinite difference, and many more complications to boot.