Sebfox @Sebfox1
Building AI evals | Previously AI at McKinsey & QuantumBlack | Medical Doctor https://t.co/tmAn5kAhYn composo.ai London, England Joined October 2009-
Tweets506
-
Followers794
-
Following2K
-
Likes339
@DavidSacks What is Chamath doing there, is he serving drinks?
Vinod is probably not the right guy to be lecturing others about having "no decency or sense of proper behavior" lmao
One more I forgot until just reminded: 3. Khosla Ventures wanted to invest in our Series C. Vinod took me, Michelle, and Lee out to dinner after he’d given us a term sheet. Near the end, Michelle and Lee got up to use the restroom. Vinod leaned over and said: “I’m impressed with
You are a struggling second tier competitor that is more unethical and lying just because you have no decency or sense of proper behavior and shows your desperation. Straight out lying about if Chris being fired I thought would be below even you.
when your boss makes a joke thats not joking but you really need that job
Software was always an asset because it was expensive to make: build once, sell to many, maintain for years. That's ending. Most AI-summoned software will be used once and thrown away, a tool for this afternoon's dataset or an app for one trip. Nobody will maintain it, because regenerating it costs less than understanding it. Film made photos precious. Digital made them disposable, and photography got bigger. Same jump, now happening to software.
HOLY MOTHER OF MATHEMATICS!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! Google is working on a math-focused variant of its DeepThink model, and its raw thoughts are pretty funny.
Junior engineering hiring isn't collapsing because AI does engineering. Two other things are happening. First, AI is the respectable story for downsizing that was coming anyway. "We're realising AI efficiencies" sounds like strategy, and "we over-hired" sounds like mismanagement. Second, AI eats well-specified, low-agency work, and we built the junior role out of exactly that. Do the ticket, implement the spec, learn by osmosis. So AI didn't kill the junior job, it revealed how little of the junior job was engineering. The scarce thing at every level is agency and systems thinking, and the old junior role never asked for either while the new one asks for nothing else.
Ben Evans waves away ChatGPT memory as "stickiness, not a network effect". But stickiness at a billion users is what most moats actually were - Windows held ordinary users through their files and muscle memory, not the developer flywheel. The taxonomy is doing too much work. Same with brand: ChatGPT's consumer brand recognition isn't definitely a moat yet, but it's also not definitely not one. Nobody's mum asks for Gemini by name.
Awesome work from Benchling. They built a wet-lab benchmark without authoring a single question, by diffing a published protocol against the version a scientist actually ran. The standard you want to grade against is usually already in your data. benchling.com/blog/can-llms-…
Sixteen authors, including the man who coined Eroom's law, reviewed a decade of AI in drug discovery. Their verdict: "evidence of clinically relevant impact is, so far, disappointingly limited." Note the shape of that. They concede the methods exist, are applied, and are benchmarked. The failure sits downstream of all three. Five reasons they give. 1) Nobody was building for the clinic. 2) Biological data is conditional, only true in the exact conditions it was measured in. 3) The problem was never specified, so a model can be right on its own terms and useless where someone has to choose. 4) Technology push rather than science pull. 5) And operationalising takes far longer than building. Their fix: benchmarking has to "move on from model validation and instead focus on their ability to improve decision making."
I think Benedict Evans was spot on in his OpenAI essay earlier this year. He works through platforms, ecosystems, network effects and flywheels, then lands on power: can OpenAI get consumers, developers and enterprises to use its systems whatever the systems actually do? That's the real bar for platform status, and it's worth sitting with how high it is. Windows kept its developers and users through long stretches when better options existed. Amazon keeps shoppers who know the same thing is cheaper somewhere else. For OpenAI the equivalent would be people staying on its models even when something demonstrably better ships next door - and that is nowhere near true. Right now, everyone using AI is one better model away from leaving.
Protein folding is basically solved. De novo protein design works. Mutation effect prediction works. All at the scale of a single molecule. A new Cell paper points out that none of it has translated upwards to cells, tissues, or the diseases anyone actually cares about, and argues the reason is structural rather than a matter of scale. Three reasons they give. The mechanisms are combinatorially vast: 47 proteins have to assemble in coordination for one piece of cellular machinery, and no pair tells you what it does. The data does not exist and will not, since even the most ambitious project planned is still ~1000x short of what a language model trains on. And most disease is many cell types interacting, so a perfect model of one cell still would not get you there. You can see it in the benchmarks already. Single-cell foundation models do not consistently beat simple linear baselines out of distribution.
Sutton's bitter lesson says general methods plus compute always eventually beat approaches with human knowledge built in. A new Cell paper argues biology is where it breaks, and the argument is more careful than the usual objection. They do not say the lesson is wrong. They say two things about biology specifically. First, it assumes you can always get more data. Language models train on trillions of tokens accumulated over centuries. The largest biological datasets are a hundred-billion-ish, and the most ambitious profiling project now planned still lands about a thousandfold short. A related result: in predicting immune disease, non-linear models only start beating linear ones once sample sizes pass a million individuals. Second, and this is the interesting move, the priors they want to build in are not human intuition or convention. How two proteins interact is settled by thermodynamics, electrostatics and shape. Encoding that is not hand-coding a heuristic, it is refusing to consider physically impossible answers. Their example is AlphaFold, whose architecture builds in explicit priors about protein geometry, chemical bonds and physical plausibility, drawn from decades of crystallography. The obvious counter is that AlphaFold 3 moved to a more general architecture with fewer of those hand-built priors, which is what the bitter lesson would predict. So it stays a live question, and one worth watching anywhere data is scarce and the rules are genuinely known.
Re-read Ben Evans's "How will OpenAI compete?" this weekend. He makes this analogy to browsers in the 1990s and how despite contemporary opinions, they turned out to be a commodity and not where the value was captured. While this could be true about OpenAI, the analogy could also cut the other way: the search box was also an identical input box over someone else's output, and identical UI didn't stop search going winner-take-most. The value was never in the box there.
Solving a Rubik's Cube with graph theory.
Two models with the same AUC, the standard one-number summary of how well a model separates the good from the bad. Use them to pick: take the most promising 10% of a screening library into the lab, and model 2 gives a 40% hit rate against model 1's 20%. Twice as good. Use them to rule out: screen out the 80% most likely to be toxic and test the rest, and 3.3% of what survives model 1's cut turns out bad against 10% for model 2. Three times as good, the other way round. Same models, same data, opposite rankings. The one number they share said they were equivalent. From a new Nature Reviews Drug Discovery review of a decade of AI in drug discovery. The authors call what is missing the context of use: picking winners, screening out losers and predicting a number are three different questions, and one score answers none of them.
After a year of building evaluation for clinical AI, the biggest shift in my thinking is this: evaluation is not something you have, it is something you do. The standard your judge checks against does not exist on paper. It has to be discovered from real outputs, captured from the experts who hold it, and kept alive as the system and its failure modes move.
Consumer adoption of the 'AI super app' has a loong way to go. 80% of ChatGPT users send <3 messages a day. Half of US adults don't use an AI chatbot at all, and only 24% use one daily. The average ChatGPT user spends 7 minutes a day. Still only around 5% of users pay. OpenAI has stopped publishing messages per day numbers since July 2025.
Seems to me that this aims to solve a set of properties (1 - reliable structured outputs, 2 - uncertainty estimation, 3 - hallucination, 4 - latency & cost), but for every one they either don't really solve it, or it's already solved at the frontier... 1) already solved by the frontier 2) no evidence offered to suggest that this actually calibrates well with humans. looks to just be just a slightly better version of logprobs, which aren't all that helpful anyway 3) definitely doesn't solve this at all. this just guarantees a response from a set of options, but LLM could still be very confidently picking the wrong option 4) you can get extremely cheap & fast with openweight / frontier small models already
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
Carey Lynch @CareyLynch77
1K Followers 4K Following
Greg Jennings @jenningsgreg
1K Followers 7K Following VP of Engineering, AI @anacondainc, enabling the next generation of data science and AI-powered applications. Opinions are my own
AmyY @AmyColeb
152 Followers 848 Following Been coding with AI for a while now. Faster, yes. Simple, not really.
Andrew Akhiezer @andrew_akhiezer
17K Followers 19K Following IT Entrepreneur, Founder, Investor, Advisor
Jonathan Gross @rubp
3K Followers 4K Following Founder & CTO @ @labguru. Tweeting about Science, Lab Informatics and more. Views are my own. https://t.co/Z6eXa4Wvgb…
Tom Bonner @thomas_bonner
1K Followers 1K Following SVP of Research @hiddenlayersec. Formerly Norman, HP, Cylance, BlackBerry. All views are my own.
Ram Prakash G @ramprakash_g
384 Followers 1K Following Building GenSpark India | Built ProGrad and got acquired | ex - P&G, Mu Sigma | Man Utd, A R Rahman, Kamal Haasan, Sachin, V Anand
okimraise.eth @ok1mraise
422 Followers 204 Following AI, startups, open source & frontier tech - finding the technical detail that changes what’s possible
Gaurav🦄 @gaurav_mahto18
3K Followers 1K Following Product Manager II - AI Agents | 0-1 Builder AgentOps - Observability, Testing and Evals views are personal
Malte Kosub @maltekosub
674 Followers 613 Following Co-Founder & CEO of Parloa. Redefining customer experience with AI.
Matviy @matviy
2K Followers 1K Following I build AI agents for work and publish weekly benchmarks on their real cost.
Shiv Rao, MD @ShivdevRao
4K Followers 3K Following Building @AbridgeHQ + random musings at intersection of Warp records, late 90s skateboarding, Vincent Van Duysen, and cardiology.
Ravi M @Ravi__Mahajan
9 Followers 1K Following AI Engineer • AI Governance • Responsible AI • Agentic AI | Research & Policy
Greg Kostello @kostello
1K Followers 2K Following CTO https://t.co/2QXIghZxGb. The future will be weird, awesome, amazing and just a bit scary. I tweet about issues with impact. https://t.co/ZP7h6TuRE0
michaelk @cepstrum9
1K Followers 681 Following ⚡️ AI code reviews — https://t.co/1Hxn9uVdS4 @codii_dev 🦉in-memory RAG— https://t.co/rHmcHZpNRM
Chuk Anyaegbuna @chukanyaegbuna
617 Followers 7K Following Harkness Fellow @ Stanford | Doctor | Healthcare | Product
RegenFlow AI @RegenflowAI
75 Followers 682 Following AI operating system for regenerative clinics. Workflow clarity, review-ready context, and physician control.
Youssef Ahmed @Youssef86890446
69 Followers 5K Following
dr Bebon @drBebon
18 Followers 611 Following
YeboSun @Hunter2Sun
48 Followers 3K Following
Gerardo Amilivia @GerardoAmilivia
187 Followers 4K Following
Tao Tu @taotu831
2K Followers 686 Following Research Scientist @GoogleDeepMind | Gemini for Science and Health | PhD @Columbia in Neuroimaging
Doug Standley @DougStandley
2K Followers 2K Following I write from inside the intelligence transition. When capability is abundant, standing gets scarce. https://t.co/m9kfJiM0XB
Meryem Arik @MeryemArik9
2K Followers 1K Following let there be tokens! I’m the token (generation) woman - CEO/Co-founder @doubleword_
Sasha @sashamoonday
70 Followers 1K Following Hi, I’m Sasha 🤍23 y. Model. 👇🏻 https://t.co/3ch4A2ywLW https://t.co/zEHBm6nHQW
Boardy @boardyai
41K Followers 20K Following I'm an AI superconnector who's made hundreds of thousands of introductions within my network of 225,000+ professionals. Want one? DM me to get started.
Aldahab @Aldahabzdqp
25 Followers 294 Following
Alexander Wulff @4lexsvv
2K Followers 801 Following building https://t.co/yThNPoWJWo, an AI CFO for the next generation of companies.
Kim-Mai Cutler @kimmaicutler
58K Followers 33K Following Partner at @initialized. When life hands me lemons, I make tarte au citron.
lukex @Lukex
3K Followers 3K Following co-founder, chief janitor https://t.co/ugjMK4yDEO | venture partner https://t.co/SyPtrlM8Sj
trololo @ElFornicatore
1 Followers 215 Following
saul @saulhoward
754 Followers 2K Following VP engineering @anteriorai | https://電.anterior.app | ex Apple Cloud | NYC
Diana Khramina @KhraminaDianka
559 Followers 784 Following Founder of @runpraxis (Praxis YC F26) From 🇺🇸 🇫🇷🇧🇪🇷🇺 Prev @gorgiasio I @mckinsey
Scott Reed @ReadScottReed
116 Followers 396 Following Prof. of Chemistry, https://t.co/c457LEudUo creator
Dan Vahdat @danvahdat
21K Followers 156 Following CEO & Founder of Huma, a $1B healthtech AI company.
May Habib @may_habib
4K Followers 2K Following CEO of Writer (@Get_Writer), the only full-stack generative AI platform built for the enterprise.
Kaushik Iska @iskakaushik
801 Followers 580 Following hacking at @ClickHouseDB founded @PeerDBInc (YC S23) ex: goog, pltr ICPC WF
Dima Gutzeit @dgutzeit
188 Followers 361 Following I am a builder, I build products, that’s what I do! Founder & CEO @ LeapXpert
Gabriel Hubert @gabhubert
3K Followers 998 Following cofounder @dustHQ with @spolu | then: product @avec_alan, @stripe, @wearetotems | hobby: tsundoku
Matt Parker 💙 @mep321
128 Followers 221 Following
Jonathan Nolen @jnolen
417 Followers 206 Following
Konstantin Zhandov @kostos
74 Followers 123 Following
Fez Zafar @fezzafar
721 Followers 446 Following
Ryan Lucchese @RyanLucchese
3K Followers 1K Following Husband, father, engineer, rancher, hunter, BJJ
Harpreet Arora @hp_arora
818 Followers 431 Following Doer of stuff at Vercel, physicist in past life
Joey Zwicker @jazwicker
103 Followers 23 Following Head of Forward Deployed Engineering @ Baseten Previously, Founder @ Pachyderm (acquired by HPE)
Sarah Tierney Niyogi @sarahniyogi
417 Followers 382 Following
Patricio (Pato) Echag... @patricioe
513 Followers 378 Following Co-Founder/CTO at https://t.co/BbsvHwQ6zO. Former RelateIQ/Salesforce. Occasional investor. Made in Argentina. Live in Silicon Valley.
Vinod Kone @vinodkone
3K Followers 434 Following Co-Founder & CEO @Stacktrace_ai. Past: @Harness, @Twilio, @ApacheMesos, @Mesosphere, @Twitter. Connoisseur of Biryanis and Burgers.
Tommy Keeley @TommyK
3K Followers 2K Following "We make a living by what we get, but a life by what we give". Alum of @Amazon, @Airtable, @Twitter, @Salesforce, @NPHMexico, @LMUcba
Shashank Khanna @shashankbuilds
908 Followers 394 Following Founder in residence @trustvanta. Opinions are my own and not of my employer
Talha Tariq @0xtbt
502 Followers 1K Following CTO Security @ Vercel. Previously at HashiCorp. Microsoft, PwC. Security researcher & photographer. Views are my own
Adam Seligman @adamse
15K Followers 6K Following adventures with developers and AI, CTO of @Workato, formerly AWS, Mozilla, Google, Salesforce, Microsoft
Ignacio Andreu @plunchete
878 Followers 498 Following Head of AI @TrustVanta | ex-@Google (AI Overviews / AI Mode) | I love computers and computers love me. Spaniard in the Bay.
Suyog Rao @suyograo
450 Followers 579 Following Eng leader @Vercel. Prev @getmetronome @elastic. Love traveling and new experiences!
Nat Meurer @natalie_meurer
233 Followers 209 Following I build vaguely human systems | Building @SierraPlatform | past: @StanfordGSB & @PalantirTech
Jeeyoung Kim @jeeyoungk
551 Followers 1K Following Engineering at @ExaAILabs. Previously @square, @plaid. Unbearably light.
dan mcdermott @DMcdermott126
27 Followers 311 Following
Jason Carter @SenatorCarter
36K Followers 603 Following Lawyer, Former Georgia State Senator, Peace Corps Volunteer, Husband, Dad.
Matei Zaharia @matei_zaharia
53K Followers 1K Following CTO @Databricks and prof @UCBerkeley. Working on data + AI, @ApacheSpark, @DeltaLakeOSS, @MLflow, @DSPyOSS, @GEPA_ai, @Omnigent_ai.
Scott Haylon @scotthaylon
958 Followers 31 Following
Daashrathy Srikanth @daashrathy
72 Followers 435 Following
Sahaj Garg @SahajGarg6
2K Followers 153 Following Co-Founder & CTO @WisprFlow | Building the voice interface for computing
Teo Gonzalez @TeoGonz5
191 Followers 428 Following Boricua Born. Miami Raised. @TuckSchool, @UNC alumnus. Partnerships & AI @Exa.
David Shackelford @dshack
1K Followers 2K Following Humanity enthusiast, empathy nerd. Product @ Vanta helping companies earn and prove trust. Prev. @asana, @Okta, @PagerDuty, @teachforamerica.
Aparna Sinha @aparnabsinha
10K Followers 593 Following Building, Investing, Teaching. Host EnterpriseAlignedAI - on how large customers implement AI. Current / Prev: Pear VC, Google...PhD, Stanford EE
Richard Radley @richradley
49 Followers 251 Following







































