Yep agreed on the more cultural in group/out group dynamics being the driving force here. But I do think there is a fundamental philosophical incommensurability between benthamism on the one hand and the standard range of left/liberal/conservative views, which all have some commitment to categorical moralism, humanism, partialism, and rights, and that this is likely to accentuate tensions including (or especially) when not made explicit, eg in pop media coverage. (Not claiming that EAs are specifically benthamites to be clear, using him as a shorthand)
@panickssery@ColinJSharpe it's not the only factor but i do think it's underrated how much people's beefs, whether they realize it or not, are actually just with jeremy bentham
@TheStalwart This is the classic in this area (‘actor-network theory’), on scallops and fishermen: ics.uci.edu/~corps/phaseii…
Honestly rhymes with how you guys defamiliarize industries by asking seemingly basic questions about how stuff works and what the relevant issues are
@TheStalwart Science and technology studies (cf odd lots guest Donald Mackenzie) / sociology of science. I’m surprised you haven’t gone down this rabbit hole, complements McLuhan and media studies well. Latour is interested in the ‘agency’ of non-human actors
@TheStalwart Striking how STS / Bruno Latour fans have been training for this moment for decades but fell over at the first hurdle due to fighting the last battle (crypto) + repressed small-c conservatism and small-h humanism
Here are the questions that currently seem most important to me:
-- Better methods for "mind-reading" model activations. These methods have advanced a lot recently. We now have multiple techniques now for decoding activations into somewhat readable language! But all the existing techniques have obvious limitations. NLAs are often hallucinatory, Jacobian lens / related methods are limited to bag-of-words readouts (and capture only part of the full activation vector). Making progress on these failure modes seems pretty tractable. Moreover, the paradigm of decoding (single-token, single-layer) activations may be inherently limiting -- perhaps we should be building techniques to decode whole context's worth of activations (after all, that's what the model uses!). There's also important work to do in characterizing simple baselines -- "just ask the model what it's thinking about" is quite powerful, and worth studying in it's own right!
-- Better methods for answering "why" questions. We're much better at "mind reading" than we are at demonstrating causal claims about what caused the model to do something. For example, we can often tell that the model is aware of being evaluated, but have comparatively much greater difficulty telling whether this awareness is influencing this behavior. There are a few directions here: (1) black-box techniques for inferring causality -- such as resampling model responses after making edits to the prompt / context -- are very powerful, but also open-ended and require some taste. Getting LLMs to do these experiments well, in an automated fashion, would be a big unlock. (2) the science of activation steering is pretty immature, and typical practices (adding a constant vector at all token positions) are pretty janky / tend to brain-damage the model. More surgical / targeted steering, or fancier methods (examples: "on-manifold" steering using activation diffusion models) could help.
-- Fitting good linear probes for unverbalized motivations / awareness. Take deception as an example -- what's the best way to fit a probe that will generalize to covert deception? Fit it on CoT excerpts where the model's talking about its plans to be deceptive? Fit it on the part of its response where it's actually doing the deceiving? Is it important that you elicit on-policy deception examples and use those to fit your probe, or can you use synthetically written off-policy demonstrations of deception? Is it important that your probe be causally meaningful / predictive of an upcoming intent to deceive, or is it fine for it to merely recognize deception in the transcript post-hoc? The same questions apply to many other concepts of interest that we'd like to probe for -- evaluation awareness, grader exploitation, etc.
-- Understanding generalization in training. The literature is now replete with "weird generalization effects" -- emergent misalignment being a canonical example. In general, when you train a model to do X, it usually learns X, and sometimes it generalizes to Y and Z. But other times it just learns X. Why? Nobody really knows! Step 1 here is probably a much more thorough characterization of the behavioral empirics -- gathering data about which X's generalize to which Y's and Z's, for which kinds of training data / algorithms. Once we have a lay of the land, we can start to connect these observations to model internals, and ideally develop tools to predict such generalization a priori.
-- Model "psychology" and "biology." The above questions are largely methodological -- building better tools to answer questions about the model. But I think interpretability research has largely underinvested in the part where you actually then go answer the questions! Currently this feels like 5% of the field, and I think it should be 50%.
Brain dump of "psychology" questions: How well can LLMs introspect? How coherent are their belief or value sets? What's up with personas -- are LLMs best understood as "writing about a character," or have they "become" the character in some sense? Does the LLM have an agenda above and beyond what the Assistant wants? Do LLMs have explicit representations of goals or preferences? How can we tell which parts of the LLM's activations "belong" to the Assistant? Which kinds of reasoning necessarily route through "verbalizable" representations, and which don't? Do models think internally in phrases / sentences, like an inner monologue, or is it more like a jumble of concepts? Is the "global workspace" claim for real -- do models have a more "conscious" part of there activations and a more "unconscious" part? Can models tell when they're on-policy, and does it matter?
Brain dump of "biology" questions: Why do models like to represent information on some tokens but not others? How do models bind thoughts or attributes to particular entities (special case of interest: how are thoughts bound to the Assistant). To what extent are important functions localized to small numbers of MLP neurons or attention heads? What parts of a model change most during post-training? Do models represent some information in a fundamentally "cross-token" way? Some concepts appear to be represented on low-dimensional nonlinear manifolds -- is there a taxonomy of these manifolds, and is their geometric structure important? Are models able to represent information in a compositional / hierarchical way, with a "grammar" that can't be described as linear / additive combinations of primitive concepts?
Dostoyevsky (The Idiot): “What is there for me in all this beauty, when at each minute, each second, I’m now compelled to be aware that even this tiny fruit fly buzzing around me in the sunbeam now, even it is a participant in all this feast and chorus, knows its place, loves it and is happy, while I alone am an outcast, and it’s only because of my cowardice that I’ve been unwilling to realize that before now!”
Everyone is doing terrible things to this poor fruit fly, trapping it in black mirror nightmare environments, inflicting max pain, forcing it to play the same beat saber song indefinitely etc, so I'm building a sim where it just gets to fly around forever in fruit fly heaven
@BronsonSchoen@_arohan_ We’re scaling up talkie and are very interested in people’s input on what kinds of posttraining and evals would be most interesting and useful
@gleech@JeffLadish@NeelNanda5@AISecurityInst you can prompt it not to think and then resample until reasoning_tokens = 0. Looks like Neel also did ICL. When I tried this with chess, for my particular prompt, Astra tried to think 23% of the time.
How much reasoning can GPT-6 Astra do without CoT? @AISecurityInst puts its no-CoT time horizon at ~30mins on math. We run the Think Fast benchmark to measure this on a wider range of tasks. Our *rough estimate* is that Astra’s 50% TH is in [8mins, 1 hour].
Astra appears to be the first model that, playing chess with thinking disabled (one forward pass per move), regularly beats the weakest Stockfish setting.
Tracks with the big jump in no-CoT math performance in the Astra system card (p. 71).
cf @RyanGreenblatt's work on safety concerns around opaque reasoning ability.
158 Followers 883 FollowingNational park cartographer and long-term investor. I map uncharted terrain and chart financial futures. My simple map: DCA into $VOO every two weeks, no matter
126 Followers 921 FollowingAnalyzing Dividends through a systematic lens. Current holdings: $VYM, $SCHD, $VIG. High frequency strategist. Enjoying watch collecting.
1K Followers 3K FollowingTrader Market watcher Classic cars, fishing & good coffee 60+ years young Always studying stocks and looking for the next opportunity
2K Followers 1K FollowingProf: Public Affairs & Pop Health @UWMadison; Author: @thegenomefactor;
From @CityofAthensPR; Education: @UTKnoxville. views my own
2K Followers 2K FollowingOnly the Machine God can save us now.
Lead Software Engineer - FinTech.
Former Spook Contractor, @microsoft, Nuclear Weapons Engineer, MQ-1 Analyst.
35K Followers 2K FollowingEnergy Abundance, AI Power & Data Centers, NBA, & Rap posts. Background in equity & credit markets, startups, & IB. Personal views NOT employers. NOT advice.
84K Followers 3K FollowingScientist at Tufts University; my lab studies anatomical and behavioral decision-making at multiple scales of biological, artificial, and hybrid systems.
18K Followers 483 FollowingBDFL @ https://t.co/MXnNmnVajl
I break software for a living, giving back by making LLMs break em less. Lifelong ring-0 resident, LA57 fan, drew a ▲ with dx12 once.
11K Followers 8K FollowingPhD candidate @CUHistoryDept dissertating on the political economy of Wall Street in the 1980s. Words @washingtonpost @theprospect @jewishcurrents
5K Followers 200 FollowingMicrocaps, jr. resources, out-of-favor cyclicals. 75% investor 25% speculator. Not FA. May buy or sell at any time. Navy vet, homeschool dad.
181K Followers 2K FollowingCo-Founder, Novara Media. Dad. Husband. Author of ‘Fully Automated Luxury Communism’. Column at UnHerd. More soon with Verso Books.
9K Followers 1K FollowingWhatever is achieved is not final; whatever we call fulfillment is a description from inside one form of life, not an endpoint for all forms of life.
480K Followers 849 Following"The Writing Guy" | Christian | Host: How I Write Podcast | I tweet about writing, beauty, and architecture | My writing: https://t.co/SOE9HtxXdi
6K Followers 3K FollowingUsing #AI and #NLP to study storytelling at McGillU. Director of .txtlab and author of the new book, Why You Should Read More Fiction.