Posts

20. A Sketch of Helpfulness Theory With Equivocal Principals

You may rightly wonder what else I’m up to, if it’s not just blogging and the occasional adventure. The short answer is that I spend most of my time and mental energy taking part in a summer research fellowship in AI safety called PIBBSS, where my research direction has to do with questions like “how does an agent help a principal, even if the agent doesn’t know what the principal wants?” and “how much harder is it for the principal to teach the agent what they want, and how much worse are the outcomes, if the principal doesn’t know for sure what they want?”. The long(ish) answer is this post. I’ll avoid jargon, mathematical notation, and recommendations to read other long papers, but some amount of that will be unavoidable. Here’s the basic setup - called a “Markov decision process”, sometimes with adjectives like “decentralized” or “partially observable” - which forms the ground assumption for the entire subfield my research direction lies in, called “inverse reinforcement learning”:...

19. Notes On Hyperbolic Blue Paint

Image
(Read post 3, “Secret Colors, Impossible Colors”, first, or this will probably not make as much sense.) You know how sometimes you set out to do something and you psych yourself up to put a lot of effort in to have to figure out precisely how to achieve some desired goal or effect, and then the very first thing you try approximately works? Yeah, that’s about what happened when I set out to mix my own hyperbolic blue paint. It turned out to be such a straightforward effect to achieve under fairly permissive lighting conditions that I’ve shown it off to easily dozens of people, and easy enough to make that I led a small group in making some in person at a small event in March 2025. For the sake of posterity, I’ll provide a short recipe here. You will need the following: Synthetic ultramarine pigment (~$1/10 g) “Blue Lit” phosphorescent pigment, by Stuart Semple (~$20/50 g) A paint base, like linseed oil or acrylic base (Optional) Kaolin powder A small scale A mixing vessel, anything from...

18. Positive Feedback is More Efficient Than Negative Feedback - A Geometric Approach

(Epistemic status: a hastily-written recreation of something I remember reading once and can’t find anymore; if you can find it let me know. The original post had some good images. Probably best read after post 11, “Why the First “High Dimension” is Six or Maybe Five”.) Positive reinforcement is just plain more efficient and effective than negative reinforcement, if both are feasible, and I can pretty straightforwardly present a model that argues strongly for it. Consider the following setup: we’re trying to get some tiny simple agent to navigate to a goal area within some simple space. No overly complex obstacles, no particular hazards, just a tiny simple agent capable of approaching or avoiding marked areas and a goal with a strong but extremely short-range attractiveness. We have negative feedback markers, which the agent will strive to avoid, and positive feedback markers, which the agent will try to approach; we can model these both as having some infinite-range repulsive or attra...

17. Schelling Points, Flagpoles, Drip-Trays

Communities exist. This much we can hopefully agree on. Some communities arise naturally, while others are founded for purpose - anything from a desire for companionship to the need for more than just a few people to participate in some activity to the assembly of a lever to move the world with. One type of community is what I’d term a “flagpole”: a community that can be either intentional or accidental, but is always about something in particular: some cause or trait or practice. It’s generally the only one of its kind in its catchment basin, whatever that basin might be - walking distance, easy driving distance, or even nested within some nebulously defined internet subculture; if it isn’t the only one, there sure aren’t that many of them. It must also make itself prominent, often explicitly advertising to a large potential audience. After all, if a community isn’t growing, it’s shrinking - so goes the aphorism. This makes the flagpole a natural place for people who wave its metaphor...

16. Maybe Big Someday, Definitely Good Now

The world finds itself in peril, on the brink of self-immolating calamity in any of a handful of ways. Bit by bit, geopolitics has been turned into a multipolar tinderbox; drive for profit at all costs and for ruthless corporate growth are an ever-hungry furnace, with the continued impoverishment of billions; the climate warms, and forests burst into flames as the next zoonotic plague looms. And this is to say nothing of the neglected funeral pyres of nuclear nonproliferation, poor distribution of food and medical goods, and any of half a dozen simmering genocides at some or other great power’s behest, to name just a few ongoing failures. With all these cause areas screaming for resources with which to fight the flames, what are we to do, wishing to do the most good as efficiently as possible, if we hope to douse the rising flames our global civilization sleeps fitfully among? The prospect of near-term AGI has only stoked the flames of our perdition, and threatens to flash into a sudde...

15. “Too Stupid to Work” is Too Stupid to Work

(You might want to refer briefly back to post 4, “Seven-ish Words from My Thought-Language”, for the concept of [vanilla-obvious]ness.) I used to find myself repeatedly running into the problem of neglecting to do something of real value because I thought it might be too obvious or stupid to do. I still sometimes do, mind, I just used to have this problem way worse. The problem is in the heuristic that something can just plain be too stupid to work, which is, itself, much too stupid to work and in fact often fails badly. To that end, let me quickly describe a few things that I think solidly fall into the reference class of “things that feel too dumb to work, but totally work”. First off, write things down if you want to remember them much later on. This can be anything - conversations you have with people, notes to yourself, shower thoughts, events that have happened that you were a part of. If you want to remember anything more than broad strokes in a year’s time, write explicit notes...

14. On Microtonal Music (Extra Bits)

Here are two extra bits to talk about that were out of scope for my previous post about microtonal music, but which I still think are valuable to read. First, some music I recommend, if you want to try listening to microtonal/xenharmonic/untwelvish music. If you like guitar-driven rock, Brendan Byrnes is a surprisingly prolific musician who’s done a lot with sharp-fifth tunings like 22edo and 27edo; I recommend starting with the albums “Realism” and “Holocene Dream”. (“Neutral Paradise” is one of my favorites, and “Micropangaea” is his older and most varied album.) For a better-known option, King Gizzard and the Lizard Wizard has done plenty of music in 24edo, mostly emulating Anatolian, traditional Arabic, and Jewish scales; “Flying Microtonal Banana” is the starring example among albums, and the two-part album “K. G.”/”L. W.” presents another excellent example. For jazz, go looking for the sweetened thirds and flat fifths of 31edo; Hear Between the Lines has published some excellent ...