75 Comments
User's avatar
Kenny Easwaran's avatar

I feel like I've been gesturing at this point in my academic decision theory papers about FDT.

Here's the recent (open-access) one, with Ben Levinstein and Ted Shear: https://link.springer.com/article/10.1007/s11238-025-10080-w

Here's the older one that probably isn't open access (but the pdf should be available on my website, at kennyeaswaran.org/research): https://www.jstor.org/stable/48692220

EDT and CDT both think that the fundamental thing to evaluate is actions, while FDT thinks the fundamental thing to evaluate is algorithms for decision-making (or as Leonard Savage would have said, the one "grand" decision of what to do in every single contingency in every possibility that might come up, the full plan for one's life). CDT and FDT both think that we evaluate these things causally, by seeing what would be brought about from actually implementing the thing, while EDT says to evaluate it by what it is correlated with.

I think that CDT and FDT are both right, but address different questions. There's one question of what the rational act is, and one question of what the rational decision algorithm is. We naively think that these two should agree, in that whatever the rational decision algorithm is, and whatever act it outputs, must be whatever the rational act is. But Newcomb problems just show that this is wrong.

The rational decision algorithm to choose doesn't output the rational act. But that's just familiar from the conflict between act and rule utilitarianism.

Elias Schmied's avatar

Hi Kenny! Thanks for the comment. Weirdly enough, I actually ultimately disagree. As you know, proponents of FDT/UDT are doing something much stranger, in that they actually claim that the correct act is the locally irrational one. I think that they are correct, but that they just don't explain properly why this is the case. (My most recent post goes into much more detail on my view of the sociology of the situation, I'd be interested what you think). My personal opinion is that the study of decision theory inevitably leads to metaphysics - that it doesn't really make sense to talk about what is "rational" without getting into metaphysics.

Schneeaffe's avatar

Two notes on Bomb:

>So for the Right box commitment, there’d actually be a one in a trillion trillion chance of getting an empty Left. Although I would have committed to taking the Right box in that case, so it wouldn’t matter.

Wrong? Omega only predicts your behaviour in the bomb-visible case, you need not commit to anything in the other. Doubly surprising because "transparent Newcombs" formulations and discussions are often confused over what exactly is predicted, and the Bomb descriptions specifically fixes that.

I also wonder if this is the best example. It involves believing that a 1/epsilon event happened, *only* if you make a certain choice, which is strange at best. Why did you chose it over counterfactual mugging?

On the theory itself: Like probably many LWers, I have thought of some version of the Great Committment, and I think it would only make you updateless about things you learn after - the fully updateless version would seem worse to your current decisiontheory beforehand. In some ways this is good, it avoids truely insane stuff like "controlling whether you exist" (link), but in many ways it also doesnt go far enough, and it sure isnt theoretically elegant.

https://www.lesswrong.com/posts/4MYYr8YmN2fonASCi/you-re-in-newcomb-s-box

Elias Schmied's avatar

oh, you’re right of course. fixing that.

Schneeaffe's avatar

I agree that the practical impact is limited in the modest reading, but its somewhat important as a theoretical difference. Many potential counterfactual muggings are blocked by this, and other ideas that you should act mostly for Tegmark worlds youre not in.

Bz Bz Bz's avatar

I think this is a good contribution to the debate

A couple points:

1) It might be psychologically possible for me to make “the great commitment”, but it’s not as easy as flipping a switch and it’s not clear that it would really be the most productive use of my time and effort. Humans seem to function fine without having to make crazy pre-commitments, so the upside is unclear, and there is also the risk of making a mistake and locking in bad pre-commitments. Future AIs with better intelligence and ability to self-modify seem like they’ll be in a better position to worry about this stuff.

2) It actually might make a big difference to the future whether we make the commitments in a CDT way or an EDT way. EDT might suggest looking for ways to acausally cooperate with distant civilizations (Evidential Cooperation in Large Worlds), whereas CDT (and the commitments CDT recommends making) doesn’t. So the traditional academic philosophical debate might actually be more important than the LessWrong discourse

Elias Schmied's avatar

Thanks Bz!

1) Since it’s a “moment of understanding” type thing, I think it might be as easy as flipping a switch!

The thing is, I purposely emphasized the vagueness of it so that there wouldn’t be worries of “oh, did I commit myself to something crazy” later - the correct notion of updatelessness is left completely open to further investigation! (and I agree it’s very non-obvious, cf the screenshots in the post). I claim that you shouldn’t be worried because if it turns out that there’s something crazy, you can just decide that you didn’t mean anything like that in the first place.

I do think there’s some practical upside in better understanding what we do on a day-to-day basis. cf “the hostile telepaths problem” and “newcomblike problems are the norm”. But of course I agree that it’s not a huge practical concern.

2) CDT definitely would if you use logical counterfactuals! I’ve kinda stopped thinking EDT vs CDT as a hugely important live debate ever since coming across Abram Demski’s EDT=CDT thing.

hypnosifl's avatar

Nice piece! It occurred to me a while ago that the claims that FDT gives different recommendations from EDT in cases like XOR blackmail could be avoided if we are thinking about a special case of EDT where we assume there are many diverging copies of you (as in the MWI, or if you are a simulation and a bunch of copies are made with small random perturbations), and you have a utility function that treats all your copies equally and tries to maximize average utility for all of them. This is equivalent to thinking about what pre-commitments you'd want to make before all the copies diverge in order to maximize expected utility. It also occurred to me later that the FDT recommendation to cooperate with other FDT agents (not divergent copies of you) in a one-shot prisoner's dilemma could be justified if we imagine that all pre-commitments are made behind a veil of ignorance about which agent we'll be, so that if there's a combination of choices both can make that allow them each to do better than the Nash equilibrium (i.e. the Pareto optimal choices differ from the Nash equilibrium, as in the prisoner's dilemma where both are better off if they cooperate than if they both defect, but the Nash equilibrium is that both defect), all such agents should pre-commit to choosing that.

Do you agree with seeing "the great commitment" as potentially a special case of EDT, or something different? Either way, do you think "EDT + imagining the most desirable pre-commitments from behind the veil of ignorance and sticking to those" would ever differ in its recommendations from FDT on any specific problems?

Elias Schmied's avatar

Yes, the thing you're thinking of is roughly EDT with alternate realities / updateless EDT! It is also my actually preferred framing. Seems like we converged on something similar :)

Though updateless vs updateful is generally taken to be an orthogonal axis to EDT vs CDT (this is Christiano's 2x2 if you want to ask an LLM about it), so commitment theory (which is basically just a different framing for updatelessness) wouldn't be a special case of EDT.

In terms of whether it ever differs from FDT, FDT is more of an umbrella term (the MIRI/OP exchange on decision theory would be a good thing to read here), so it's hard to say.

Raj's avatar
Aug 2Edited

I think one-boxing is easy, but I don't know if I can left-box. I don't know the exact terminology that captures my intuition. Maybe normative uncertainty, but at the level of decision theories? The logic goes the exact opposite direction for newcombes: even if you think two-boxing is correct, you might still consider one-box out of a kind of epistemic humility that you might be wrong, and don't stand to lose much.

Like, decision theory is cool, precommitment is cool. But at the point that you are actually facing imminent death, wouldn't you obviously do a kind of moorean shift and be like 'screw this sophistry I don't wanna die'?

Quinurum's avatar

Ok, so, on how I think that DBDT can generate true updatelessness.

You imagine an idealized rational agent with all the information you had at the critical point. Then you ask yourself if such an agent would adopt the disposition to pay counterfactual muggers he encounters or not. (or *the one* counterfactual mugger the one time he will encounter him). Since your CP-self didn't know how the coin would fall, he adopts the disposition to pay (from his perspective it's higher EV).

This wouldn't require you to actually have considered the scenario at CP, you can just think about what disposition your idealized past self would have adopted.

Also idk if I was arguing past you a bit because I like 99% agree with you. Especially with that DBDT and commitment theory are actually orthogonal to what is rational, I think that's a really great point. I think the only differences between this DBDT version and your commitment theory is that

a) DBDT talks of dispositions instead of commitments ("adopt the disposition that is good to act out of" instead of "commit to those actions that would have been good to have committed to. I think I like this a bit more, since I feel like commitments are something more specific, like a "self-set external rule" while dispositions are just "the way you're inclined to act" (you can solve Parfit's Hitchhiker by being disposed to always keep your promises, and I think this plausibly involves no committment in the sense above. But maybe that's also just a verbal difference idk).

b) Fisher allows for those "Critical Points", and I really like that idea. I think it's probably true that our dispositions can change throughout our life (and that we can actively work on them), and the critical points adress that imo. Also I think setting an "universal critical point" at the beginning of time would actually mirror your "great commitment". And the search for where the critical points for our actual actions are would probably mirror your proposed search of which situations are, in which way, covered by the great commitment.

Elias Schmied's avatar

Yeah, that makes sense. The issue is that disposition-based decision theory still talks about many different dispositions, right? But yeah, the way I frame the great commitment as a psychological attachment at the end is kind if like a disposition, so something like that definitely works.

Quinurum's avatar

Yes, but I think if you want one overarching "great disposition" you can choose one disposition that governs any possible decision, with the CP (which is the information cutoff, basically) at the beginning of time.

But obviously it's impossible to actually adopt anything close to that (or to program a machine that does that, there's just way way way way way too many possibilities to consider) but what you can actually do is adopting different dispositions with different CP's for different decisions. (even if the "great disposition" would be more optimal)

Elias Schmied's avatar

I think it is possible though, like I argue in the post. It doesn't need to cover all the possibilities, it can be vague and the details can be filled in later. And if you do many dispositions, you'll fail on situations you didn't anticipate, like we discussed on Twitter

Quinurum's avatar

Ok wait maybe it is

The thing I thought you obviously can't do is figuring out how exactly the "great disposition" looks like and then check for every decision what the great disposition tells you to do. Because for that you'd need to consider every possible path the universe could have taken (or at least all the paths were "you" exist). And I think this is also impossible in commitment language, the great commitment will always stay somewhat vague.

But yeah you can say that you can figure out if some particular course of action you consider is part of the great disposition/commitment without knowing the whole thing (if that's what you mean).

I wonder though why evolution has equipped humans with such strong "brute CDT" intuitions in so many cases. Maybe it's just an easier/more economic way of thinking that usually works just fine?

Elias Schmied's avatar

No, that's not what I'm saying, I'm saying you can figure out whether the great commitment applies only when you actually get into a given situation.

I'm not sure if it has, many people have EDT intuitions on Newcomb's, and concepts like integrity and honor have an FDT flavor to them.

Hein de Haan's avatar

(Will comment on the post as a whole later!)

"even though we have no reason to think these other versions exist".

Huh? These other versions actually exist in a Tegmarkian multiverse, for example. Not saying that's necessarily the correct view, but "no reason" seems much too strong here.

Elias Schmied's avatar

Yes, I’m speaking to a “common sense” perspective

Julian's avatar

> You know that Omega has a failure rate of one in a trillion trillion in this situation.

I am a bit confused by the wording of this bomb scenario.

Does this failure rate mean a failure rate of prediction or a failure rate of putting the bomb or not bomb correctly into the box despite having made a correct prediction?

Elias Schmied's avatar

Of prediction!

David Piepgrass's avatar

You lost me at

>But wait - Omega was predicting me, as usual, so this is definitely covered by the great commitment.

why is it covered? Also

>LessWrong-style decision theory

is an extreme name; I doubt almost anyone from LessWrong takes the left box. You are talking about Yudkowsky's FDT (Functional Decision Theory), and while many people on LW love Yudkowsky's Sequences including myself, they don't include FDT.

Eliezer Yudkowsky's avatar

- If you let CDT make a grand commitment it turns into Son-of-CDT which is not the same as FDT and loses all precommitment races. https://www.lesswrong.com/w/son-of-cdt.

- @Easwaran: FDT does not evaluate 'algorithms' as things that can be varied by intervention. It evaluates 'outputs of rigid algorithms' as the variable whose value it itself is determining.

I'm afraid that from my perspective these are common misunderstandings of logical decision theory, not isomorphs that might be easier to explain.

Elias Schmied's avatar

Hi Eliezer, I address this in a footnote, will reproduce it here:

"A key point is that humans (and likely almost all agents in the universe when they realize that they should make the great commitment) face basically no Son-of-CDT issue - there is no one that has been predicting their pre-great-commitment self with enough fidelity for it to matter. (It would have to be high enough fidelity to distinguish between making the great commitment and adopting FDT - hard to imagine, since they are so similar)."

teucer's avatar

I'm not convinced the Great Commitment implies choosing the left box. Omega predicted what I would do *in that situation,* and the situation includes my knowledge that the left box means death. As much as I would dislike cleaning up the glitter, I am quite sure my present self, who is not facing the boxes, would precommit to not dying.

Boring Radical Centrism's avatar

I think FDT's validity could be proven by running some sort of bot, or human, tournament and seeing which type of theorists get more points.

grumboid's avatar

I like this idea, but also I have to be realistic: there's an upper bound to how much disutility I'll accept in order to fulfill a precommitment my past self would have made.

If I have to pay someone a few hundred dollars because my past self would have wanted me to: sure. Forgo a larger sum of money, when I don't technically need it to live: sure. If I have to burn to death: forget it, my past self is on his own.

I recognize this will lead to more disutility overall -- for example, in certain very contrived scenarios I'll wind up unnecessarily cleaning glitter out of my living room with high probability. I don't view this as a positive, but I'm going to have to accept it. I can't precommit that hard.

Peter Gerdes's avatar

But no the following is wrong

> So maybe we should do a kind of "generalized precommitment" for those situations, where we commit now to always doing whatever it would have been good to have precommitted to?

Because you can always imagine a situation where you are punished for doing exactly that. It's one of the great issues with FDT because it wants to say that it is superior because actually adopting that way of thinking gets a better result but that can't always be true because the situation can always be one that punishes you for doing just that.

That's why it's so important to be clear about what is being claimed. Are you analyzing what it means to be a good decision or indicating the algorithm that would be better to use or what. And once you try to be really clear it's no longer clear there is anything left for FDT to actually say.

Peter Gerdes's avatar

I mean yes, but it's not clear there is anything left once you take this approach. FDT has always been what classical deciscion theory says to do when you adopt the idealization that you first commit to an algorithm and then follow it.

The only claim of philosophical interest here was that somehow that way of thinking about the situation was somehow *required* by deciscion theory. Without that all you have left is something pretty trivial.

Elias Schmied's avatar

FYI definitely think FDT technically claims something more than this, e.g if you check out my most recent post. It’s just that it doesn’t matter much in practice.