Some news: 5 months ago I joined @coreauto, a small research lab focused on new ways for models to learn. I’m also hiring four interns to do something my younger self would've loved.
It started with a pretty unusual project. We set up a handful of small online businesses, and
Deep learning is still not solved.
It has been great to see the creativity of the community and we all have more work to do to understand the universe.
Thank you for being in this journey together ❤️
Over 15,000 submissions to One Layer Deeper. Nobody solved the problem as intended, but there are some interesting ideas for learning reusable operations and composing them into deeper computations
coreauto.com/blog/what-the-…
Project Wallfacer: I am going to stop vague posting about better than gradient descent because 10k agents, 88hrs and 300k gb300s can figure it out. Note: I would be okay if that happens and result is public, otherwise feels too powerful tech.
We should take it seriously.
The main limitation is that there are actually very few neolabs putting serious compute into fundamental research instead of doing rlaas or open weight transformers, but there are few who are doing it, are determined and moving fast.
How seriously should we take the possibility of one of the neolabs making some wild algorithmic breakthrough that puts it ahead of OpenAI and Anthropic?
How to live your life:
Meeting Mondays
Thinking Tuesdays
Work Wednesdays
Trying Hard THursdays
Automation Fridays
Sleeping Saurdays
CUDA Sundays
And repeat!
AI discourse often focuses on scaling because of eye popping numbers spent on datacenter compute. Credit assignment is very difficult
Most people from outside big labs and even many inside get this credit assignment wrong.
It’s worth doing a simple thought experiment:
It’s may 2020. GPT-3 paper just got released.
We have two diverging timelines:
a) we have the same algorithmic progress that we did since gpt3, but we cannot spend more compute on training models than was spent on gpt3
b) we keep scaling up gpt3 and we spend as much compute as we did on gpt5.6, but on gpt3 training system
Reality is that model from timeline a) beats model from timeline b) on every axis and it’s not even close
143 Followers 2K FollowingMarket memory for AI. Public research on comparable stock history — ranges, sample sizes, and the studies that fail. Free. https://t.co/iE1WwMkdY0
3 Followers 81 FollowingJust a caffeine-powered algorithm wrangler who pretends to trade and code while secretly plotting world domination—one coffee cup at a time.
8 Followers 4K FollowingDrug addict. sic semper tyrannis. Please, place me with honor at the bottom of the great pile of bodies to come. My ancestors will look down and smile.
51 Followers 487 FollowingBuild assets. Not outputs.
AI operating systems for people done rebuilding context every session.
· Reset Cost below
↓
https://t.co/3RefBBeQ7K
9.1M Followers 13 FollowingYour Only Source For Professional Dog Ratings Instagram and Facebook ➜ WeRateDogs [email protected] | nonprofit: @15outof10 ⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
2.7M Followers 3K FollowingResearch, News, and Commentary from Nature, the international science journal
For daily science news, get Nature Briefing: https://t.co/wGmQlQ8a4D