107. If We Had a Mature Theory of Agent Foundations, How Would We Know?
(Epistemic status: A vital question people don't seem to have asked themselves. A thought experiment run in public, though the credences are soft enough to spread on toast and the butter is closer to butterfly. Roughly a third of me suspects the central question below is malformed; that mass is doing honest work in every paragraph, and I won't pretend otherwise. Published while I'm at a conference: if you have errata, quibbles, or improvements for me, bring them.)
Suppose you woke up tomorrow to find, there on your desk, an elegant but battered tome: a three-hundred-page brick titled "A Theory of Agent Foundations". Suppose it fell out of the sky: the perfect solution to correctly paradigmatize agent foundations. What an agent is; what a goal is; what
values are, and what human values are and should be; how to detect all of those; measure them; taxonomize and typify and
employ them. Suppose all of that were simply there, neither requiring or offering any explanation - all that and more, three hundred pages of text with nothing more than dry frontmatter.
Here is the field-wide version of the Button Test, that old saw. You may have the mature theory of agent foundations, in full, for the asking, as long as you can specify what it is you want - more, actually, you have what looks to be the exact sort of thing you'd want, if you knew a little more and were a little surer of yourself. But thereafter you are bound to use it: no rejection, no further refinement, no second chance. Would you take it? What would have to be true of the thing, for taking it to be sane? And what goes wrong with your first naive stab or three?
I confess something of a trap in my own thought experiment. That is - the theory is stipulated correct, and by assumption I have ruled out the worst historical fate: that of epicycles and phlogiston; those things which felt like mature theories and were even predictive, and turned out to be dead wrong. But a stipulation that you can't verify from inside is worth the same paper the sky printed it on. I concede that this stipulation doesn't dissolve our present problem, but merely relocate it from "is it true" to "could we tell". There is no orbit for this metaphorical rocket to visibly reach. Nonetheless, we are adrift all the same: without a sense of what victory would look like, how can we say that we pursue a win?
Let's step back for a moment to sketch out a taxonomy of what has happened historically when human researchers have adopted a formalization of something they cared about.
- There's the ones that were propounded, and that everyone hoped for; the things that were never proved, ultimately accepted with a shrug, and haven't failed yet, where the consequences of the collapse would be earthshattering. The Church-Turing thesis is the classic example: three different people built three independent formalisms of computation, mostly not looking at each other, and ultimately ending up with three provably equivalent views of the same process. And while decades of counterexample-hunting have come home empty-handed, no one's ever proven the thesis... and yet near everyone still believes it. Conviction by convergence is a real thing that really happens, and we know it by its smell.
- There's the things that everyone used confidently before anyone understood quite why they worked. This is the province of calculus - running a century and change on handwaved infinitesimals before Weierstrass handed out both problem and solution - and thermodynamics - a theory of mounting disorder and caloric fluid predated statistical mechanics.
- Of course, there's the ones that were used confidently, and that had real predictive power... and then turned out to be dead wrong. These are your epicycles and your phlogiston. Epicycles let Ptolemy compute eclipses thousands of years out; phlogiston organized real regularities of combustion right up until it burned to the ground. Predictive success may be evidence of something, but it need not be evidence of the right bones.
- There's the sad case, the Cassandras: correct answers that bounced right off of deaf ears. Wegener had the right answer about the continents but spent fifty years in the wilderness for want of a mechanism. Mendel and his ratios went unread for thirty-five. Even Semmelweis died after over a decade of mockery, never mind how starkly right he was, nor how starkly deadly being wrong had been.
- On the other hand there's the other case: the formalization that reshapes a field around itself almost on arrival, for better and worse. Null-hypothesis significance took the field by storm, becoming the definition of scientific inference for a century and finally being grounded in philosophy of science by way of falsificationism, testifying in its own defense all the while. Homo economicus did the same for behavioral economics. The gold-foil experiment obliterated the plum pudding model and gave us the earliest version of the atom as we know it now: full of empty space. A bound theory goes beyond being simply believed, becoming infrastructure - and infrastructure grades its own homework.
- Finally, there's the weird case: those frames which dissolve without a verdict. Cybernetics is the most famous example, and the only one that comes to mind. It was the first serious run at exactly our own subject: feedback, control, purpose abstracted from both animal and machine. It was also never proven nor disproven, dissolving away into control theory, neuroscience, information theory, ecology, complex systems, and half a dozen engineering departments. As it turns out, the fate of theories is not exhausted by merely "true" (and its variations) and "false".
If you ask me which fate awaits the real thing, the next time it rolls around - not button-stupulated - then I'd spread my chips roughly like so: ~35% dead wrong, given how unpromising the base rates are; ~25% shrug-and-fine, much like Church-Turing or P ?= NP; ~25% used confidently largely on vibes and only later on grounded as the right guess; ~10% a category error that ends up provably the wrong tree to bark up; the remainder some secret fifth thing. The button version zeroes out the first worst fate by fiat. That's precisely why the button test is the easy version... but that in turn is why it's a thought experiment worth running first.
And we have run this experiment before; the tune's in a different key these days and the band is quieter for the moment but the song's the same. In the late oughts, AIXI plus VNM was that very sky-given answer: ideal induction bolted to an ideal decision theory, agency solved in principle, and nothing but engineering remaining. But it's a guess at once too general and not general enough; it demands compute that the universe cannot contain, and then immediately crashes - in several distinct and frankly silly ways - the instant you ask it to find itself inside the world (or worlds) it's modelling. Embedded agency is the nail in the coffin. It's a little shameful how long the field took to book AIXI + VNM's funeral, and all the more instructive for it: the field can, actually, evaluate a candidate... slowly. Sort of. Non-unanimously, with some parts clinging, some parts cutting, and some parts winding down the search. That's n = 1, though, and unless you count some truly old deep cuts, no one's tried to propose a second. (If you want a larger cemetery, behaviorism and cybernetics are the neighboring plots. One banned looking inside the agent and collapsed for that grave sin. The other studied the right thing, but died of quiet starvation and exhaustion.)
Anyway: enough of all that. Back to the desk: the pages are there and so is the button, bright-red and molly-guarded. Never mind you - what would I actually do? I've mulled it over a few times, and I've run that simulation enough to know what I'd do. Realistically... kick things off by staring in horror for a couple of minutes. Then get to work photographing every page. I'd send a few to trusted friends - have I gone insane? No? Worse luck - and then I'd get down to the real work. I'd read the introduction carefully, and then go looking for a main theorem, something really juicy; then read backwards from it, as is right and proper. Hunt the key definitions, which are secretly always the most important parts. Hunt the technical lemmas, because you can hide a lot of sins anywhere boring-looking but structural. Hunt the worked examples, because a theory of agency with no worked examples is a theology - and a terribly suspect one, given the sheer number of agents floating around. Look everywhere for the kind of clean conceptual engineering and novel-but-resonant philosophy that the thing would absolutely have to contain were it real. And what would make the fur on my back stand on end? Anything like:
- Any kind of clean classification of agents: something like Thurston's eight geometries, or a periodic table, or even just biology's stamp-collection;
- A procedure - a procedure! - that takes in an agent and returns its goals and values, with error bars;
- Anything about the implicit values and goals of aggregated agents, especially any impossibility results on what can aggregate and what goals they can then have;
- Anything about how agents change over time: whether and how values drift, and whether they can be steered - and how, by whom, and at what cost.
My friend and research associate PR, when I ran this past her, added two demands and a method. Namely: it had better predict the outcome of any well-designed psychology, sociology, or economics experiment you throw at it, and likewise of the outcome of any RL training run; and it had better make testable predictions - that kind that you actually test. For the record, her reading protocol: read the frontmatter, then a handful of pages sampled from deep inside, and then the entire thing, cover to cover. I'd read backwards from the theorems; she'd sample deeply; we'd both then swallow whole. Between the two of us you might even triangulate a revelation.
It seems to come down to seven checks, then, seven checks to see whether this book is a sacred text or wastepaper:
- Does it dissolve the standard confusions rather than answering them by fiat? Do the old questions come out visibly mis-posed, and can you say how? Do old boundaries vanish in new light?
- Do several independent formalization attempts, suitably posed, all land on it?
- Does it survive self-application, describing the theorist and the theory-holding process without a homunculus smuggled in under anyone's trench coat?
- Does it pass ground truth by construction? If we build an agent with known innards and hand the theory its outsides, does it recover what we hid? (This would be the closest thing to touching grass that this field is ever offered, and anything that forces our hands into actual dirt is a thing to be cherished.)
- Does it correctly retrodict the world around us? There are agents everywhere in the world around us [citation needed]. Does the theory correctly predict what they do and why?
- Is it fruitful, in the sense of Shannon? Does it let you do things that seemed impossible just the day before?
- Does it survive optimization pressure? If you aim an optimizer at it, does it deform, or did we find a True Name?
Suppse instead that such a battery of tests cannot exist, that for some reason, recognition is impossible in principle. Suppose that no amount of dirt under the fingernails could ever rationally settle whether that sky-answer is the correct answer. Then on my model, the single best route through AI to a future truly worth having is closed, permanently, and... I'm going on a very long, very extravagant vacation, is what I'm doing.
That said, two ropes let us climb back up that precipice, and I can show you exactly what they're anchored to. The former is cheap and real: verification need not yield perfect certainty to yield value, just raise the price of fooling us. Every check that we can cheaply run makes deceiving the checkers more expensive, and expense is a real quantity in this universe even when proof is not. We might well end up having to grade our (incomplete) philosophy homework with the help of the very AIs the homework is about... fine. An imperfect grader can still fail the worst problem sets, and a world where fooling us costs a fortune is still meaningfully safer than one where it's too cheap to meter. The latter is a creed, and I'll state it that way so that you can see precisely where I stand and where it's load-bearing: "everything that is knowable in principle, a human can come to know, given time". (And don't give me anything about us not having enough time. Live in your least convenient possible world, for once.) If you reject that creed, then a good deal of this post vanishes for you, and I'd genuinely like to know which paragraph went first.
Then what would send me off to check off my bucket list all the same? Some kind of diagonlization. Something Gödel-shaped, and if it comes at all I'd bet it comes out of embedded agency or its refinement: something like "no characterization of an agent is robust to that agent's learning of it", maybe, or "no characterization of agents, goals, or values covers every reasonable operationalization of them". Partial versions of this already exist - the Löbian obstacles are not - and the field's best machinery - reflective oracles and logical induction among it - is exactly the attempt to route around the partial versions. Whether the routing generalizes is, as far as I can tell, still open.
So here's your homework, and mine: go try to prove the theorem that ends this post. If it exists, I want it in hand before I've spent the decade; if it doesn't, then every honest failed attempt is a brick in the foundation of the thing we're trying to build. Press the button, don't press it, whatever: just set down, in writing, before you wake up to a mysterious textbook, exactly what you'd need to see to make your choice.
Comments
Post a Comment