A plan for a perpetual motion research center that runs forever
There is a popular tradition on this website of clarifying complex ideas through somewhat heavy handed, metaphorical dialogues. TO WIT,
Alice says “I don’t have a good lead on how to build a perpetual motion machine, but I want to find one and am rich as all hell. Here is my plan for a research center that will study perpetual motion forever that won't have any power hookups or solar panels, because it's a clever research center that does not require any external fuel”
Bob says "One of your sentences is wrong. You either have a lead on perpetual motion design worth investigating or something stupid. I lean toward stupid."
Alice says "It sounds like you think perpetual motion is impossible due to some conservation law. However, there are holes, such as around dark energy, or the fact that energy isn't conserved in all large scale general relativity geometries"
Bob says "Your plan is either a perpetual motion machine, that we should investigate as a perpetual machine design and not a research facility, or stupid. This is unrelated to conservation laws or whether perpetual motion is possible, neither of which I mentioned. Your plan will not unfold as you expect unless you designed it with a perpetual motion solution in hand, which you say that you do not have. I know without looking at your plan that if I start at its HVAC compressor and start following electrical wires, I will find something incompatible either with reality or the claims you have made. You must be claiming to expect to be surprised by a solution to perpetual motion appearing during the facility's construction instead of during its operation"
Importantly, it is logically possible for this conversation to continue:
Alice says "Well I'll power my facility with this bowling ball sized rock I found that stays at 1000C no matter how I heat or cool it."
Bob says "Holy shit, lead with that! However, while I am surprised, I wasn't wrong: why did you claim that you did not have a lead on a perpetual motion machine design? Let's, um, go get you checked over for radiation sickness, but barring that we should absolutely put resources into investigating your rock- probably in a conventionally powered lab"
===============================================================
A scalable way to produce alignment progress (that is, one that produces progress arbitrarily far beyond what the organizer could produce through their own sweat and tears), that isn't a correct alignment solution in its own right, is another such obviously impossible thing.
The analogy is between a perpetual motion research organization that runs forever, and an alignment research organization that scales. The organization founder doesn't know how to set an intelligent system to stay focused on a specific goal as its mental capabilities snowball, so he builds a system of intelligences to ponder the question, with the full intention of that system snowballing.
I was pushed toward this line of reasoning as I noticed a repeating pattern: alignment research organizations are frequently trying for venture capital or effective altruism style scalability. Field building, mentoring, bootcamps, grants, benchmarking- these are considered exciting directions because people are faced with a big problem and want unbounded impact, and I fully cede that these are far more impactful than staring at a LateX document alone and thinking really hard. However, if you still need to do alignment research you are not capable of doing so while scaling! I picture a wildfire where the firefighters are only willing to fight it, not by hand digging trenches or wetting down brush, but exclusively by starting controlled burns, because it's so big that only such drastic measures are up to the task. I note that unlike the wildfires in our world, here the fire started as a controlled burn and no one has ever put out a controlled burn before.
Narrowing down from the larger view where I'm slandering mentorship and community (slander that I unfortunately have to stand by, since they have mostly built feel-good-flavored capabilities)– Recently, a large number of people have put forth scalable plans specifically for converting money into alignment progress, and appear to be executing these plans backed by millions to billions of dollars. If you buy my argument that a scalable plan for converting money into alignment progress is an alignment scheme, then it should be judged as such: with intense scrutiny and a security mindset.
A specific instance of the problem is the recent launch of Coefficient Giving's project tailwind.[1] It is a remarkably scary instantiation of a classic plan:
Repeatedly put out calls for grant applications
Everyone choses whether to write an application, and whether to use AI (that can one-shot Navier-Stokes existence and uniqueness) to help them write.
Judge the applications by some criteria (optionally assisted by AIs that can one-shot Navier Stokes existence and uniqueness)
The winning grant recipients use the money to do alignment research, each applicant decides whether to be assisted by AIs (that can one-shot Navier-Stokes existence and uniqueness).
When viewed with clear eyes, this is obviously just a bootleg implementation of scalable oversight and I would argue it should be paused, in my impossible dreams indefinitely, or at least replaced with one of the vaguely plausible scalable oversight designs that have been examined though the lens of being risky and probably destructive. It immediately creates Amodei's country of geniuses, trying to get the money, who can think faster than you and try more approaches than you can think of. It's just that they're scaling-pilled human-AI centaurs winning mostly by parallelism instead of AIs in a datacenter winning by serial speed. Of course, if you do think science style grantmaking and publishing is already a solution to alignment (our miraculous hot rock), make that claim so we can debate that point instead of hoping that science as she is played will find a finished solution to alignment.
My object level prediction of what will go wrong with tailwind is pure extrapolation from history: when alignment research schemes seek to use ambient pools of itchy money to scale up their impact, they either fail to scale, or scale explosively while differentially producing capabilities. And then we impact-weighted average over a bunch of them. However, this is getting into the weeds of whether scaled grant funded/funding orgs are a good alignment solution, but all I'm asking for in this essay is to evaluate members of the broader class of scalable plans as alignment solutions before implementing them. I don't present ironclad evidence here on whether any specific scheme passes that evaluation. A positive prediction of what specifically will go wrong is not important compared to the diagonal argument that if you had any reason to suspect nothing would go wrong, that reason would be an alignment plan strong enough to obviate the need for the program.
Plausible plans should explicitly include why they don't scale their impact arbitrarily far past the plan-proposer’s individual effort, for the same reason that engine designs that don't claim to be perpetual motion machines always have limits to how long they can run without external inputs, limits that can be easily read off the schematics and design reasoning.
There is a tight analogy between a perpetual motion research organization that runs forever, and an alignment research organization that scales.
If I’ve gotten my point across, the answer to the following question will be clear (the thesis is a tool that answers this question among others) In my exposition, why is the reference class assistant AI used for writing grant proposals a swarm AI that cracked a Millennium Prize but costs million dollars per query, not an agent AI that can be subscribed to for $200 a month? What simple change could be made to the Coefficient program to force applicants to use the $200 AI? Why would other changes not have this effect?