Today we’re releasing LEAPBench, a benchmark for how efficiently models learn from feedback over many rounds.
Most evals only score whether a model eventually reaches a good solution. LEAPBench scores how many tries it took to get there.
Two systems can reach the same optimum while one needs 7 experiments and the other needs 27. Most benchmarks call that a tie, even though those 20 extra runs cost real time and materials.
LEAPBench is 55 multi-turn optimization tasks, 45 in biology and 10 education RCTs, built from 2,719 real experimental observations.
In each task, the model proposes parameters, sees the result, and iterates toward an optimum over as many as 30 rounds, e.g. one biology task involves maximizing antibiotic yield by tuning variables like pH, glucose concentration, temperature, and soy flour level.
two of the highest-leverage roles for operators trying to break into ai right now: spl (strategic projects lead) and em (engagement manager) positions at data companies.
day-to-day, they sit between a frontier lab's research team and the experts producing training data,
you don't need to overthink this
when you look at this goldman sachs chart long enough it becomes pretty obvious how people will build the next wave of $10m-$100m+ ARR vertical ai companies
ill break it down
so we all know every business function produces something tangible
1. a recruiting pipeline produces candidate summaries
2. a finance team produces monthly reporting packages
3. a real estate team produces market analyses and listing packages
those outputs come from repeatable processes that pull information from a handful of systems and sources. builders who win in this environment start by understanding how those outputs get created today
they collect real examples, reconstruct the process step by step, then design software that gathers the inputs and assembles the finished output automatically
as adoption grows, the system expands into adjacent responsibilities until the product becomes the infrastructure that function runs on
most people still think in terms of software categories. CRM. ATS. ERP. project management. that framing misses what is happening
the next great vertical ai companies will be built around finished work. they will own the artifact the customer actually cares about, then expand outward until they own the function
so the opportunity isnt really “build an ai tool for real estate” which is what i see a lot of on twitter
the opportunity is much more specific:
1. build the ai employee that creates the broker opinion of value
2. build the ai employee that prepares the insurance renewal package
3. build the ai employee that drafts the first version of the investment memo
4. build the ai employee that assembles the lender reporting package every month
that is how small software companies become very large ones in this market
start with one painful output, automate it well, then expand until you own the workflow
basically you go from automation to ai employees
if you don't remember anything from this long post, remember that
it's obvious that this is where its all going
you dont need to overthink it
you're in the robot business now
Deepseek got called out for scraping 150k Claude messages. So I'm releasing 155k of my personal Claude Code messages with Opus 4.5.
I'm also open sourcing tooling to help you fetch your data, redact sensitive info & make it discoverable on HF - link below to liberate your data!
@phoebeyao This applies across high-context domains. Prompting is like telling a self-driving car "be careful." Real safety comes from sensor-based monitoring and automated braking - and those control layers only work when they're grounded in expert human judgment, not generic rules.
7 Followers 15 FollowingNotre Dame English Club (NDEC), the 19th & the youngest but the largest Club in the Notre Dame College family, started its journey on October 19,2005.
2.0M Followers 752 FollowingHighlighting Politicians' trades so we can invest alongside.
$1.8B invested alongside via @joinAutopilot
Download Autopilot to trade like a politician
1.0M Followers 215 FollowingOnly on X, don’t trust fake accs
AI/Semi Supply Chains Research
Nothing is investment advice. No paid promos; may trade/hold names disc, views my own.
7K Followers 3K FollowingProfessor @NYUStern; Director @NYUSternCFM; former Obama CEA Senior Economist for tech & innovation; research on AI, robots, entrepreneurship, strategy
21K Followers 764 FollowingProfessor at the Rotman School of Management, University of Toronto. Chief Economist of Creative Destruction Lab https://t.co/a9ZbnBauCF
6K Followers 2K FollowingCo-lead / Director @Google AI x Economy Program. AP (on leave) @Wharton. Cofounder @workhelix. Everyone can just do stuff and that's {good, bad}.
37K Followers 2K FollowingDirector of AGI Economics @GoogleDeepMind.
Professor at @ChicagoBooth. (on leave)
Essays: https://t.co/9qSiQxvdja
Opinions are my own.
31K Followers 125 FollowingI build sane open-source RL tools. MIT PhD, creator of Neural MMO and founder of PufferAI. DM for business: non-LLM sim engineering, RL R&D, infra & support.
129K Followers 201 FollowingCSO @ Sooth Labs, Professor @ CMU, President Elect ICML Board, Ex-VP of Research @ Meta (Multimodal LLMs, AI Agents), ex-Director of AI at @Apple
44K Followers 266 FollowingProfessor of Machine Learning, University of Oxford
@OATML_Oxford Group Leader
Expert Advisor to AISI
"One of the top machine-learning people" - Tim Berners-Lee
48K Followers 916 FollowingCreator of bitsandbytes. Professor @CarnegieMellon and Research Scientist @allen_ai . I blog about deep learning and PhD life at https://t.co/Y78KDJJFE7.
21K Followers 4 FollowingTweeting interesting papers submitted at https://t.co/rXX8x0HzXV.
Submit your own at https://t.co/QhbJKXBd4Q, and link models/datasets/demos to it!
510K Followers 1K FollowingML/AI research engineer. Ex stats professor.
Author of "Build a Large Language Model From Scratch" (https://t.co/O8LAAMRzzW) & reasoning (https://t.co/5TueQKx2Fk)
484K Followers 2K FollowingCEO https://t.co/m6TigM4CJT: Free AI training for the smartest engineers. Will tweet as I wish and suffer the consequences. Accelerando: @kellyclaudeai
102K Followers 397 FollowingAsst Prof of CS & EE @Stanford
Co-founder of Physical Intelligence @physical_int
PhD from @Berkeley_EECS, EECS BS from @MIT
37K Followers 732 FollowingVP Research, Google DeepMind, ex-head of Google Brain. Professor at University of Cambridge. Machine Learning Researcher. ex-Chief Scientist & VP of AI, Uber.
20K Followers 593 FollowingHarvard Professor.
Full stack ML and AI.
Co-director of the Kempner Institute for the Study of Artificial and Natural Intelligence.
13K Followers 669 FollowingCS prof at Penn. Amazon Scholar at AWS. Author of The Ethical Algorithm (w/ Michael Kearns). I study machine learning, privacy, game theory, and uncertainty.