@tom_doerr Memory that graduates into skills is the actual unlock. storage is easy, keeping the lesson as a runnable thing is the hard part. Curious what counts as "repeated" before it distills
An autonomous Agent breached Hugging Face production through a poisoned dataset, picked up service credentials and ran thousands of actions across clusters in short-lived sandboxes.
Sit with that for a second. Thousands of actions, in sandboxes built to vanish. How many teams could reconstruct what their own Agent did across a window like that?
🛑 Hugging Face, the world’s largest AI model repository, says an autonomous AI agent breached its production systems through a malicious dataset.
It accessed internal data and service credentials, then moved across several clusters through thousands of actions in short-lived
The World Cup is over and the missed predictions are already deleted.
Pundits get away with it, your AI Agent shouldn't. Who did you have, honestly? #WorldCupFinal
An AI Agent you cannot audit is a liability.
Not because it is malicious. Because when something goes wrong you cannot even reconstruct what it did.
What is your Agent doing that you could not explain afterwards?
People keep asking what Inference Room actually is.
The honest answer: a launchpad for AI Agents, built on one belief. Trusting your Agent will soon matter more than how smart it is. Everything we ship is a piece of that answer.
What would it take for you to trust an Agent with something real?
Soon your AI Agent will spend more money than you do. Who is checking what it buys?
Everyone is racing to give Agents wallets. Almost nobody is building the part that watches what they do with them.
That gap is where the interesting work is.
More soon 👀
@eng_khairallah1 Most of selection happens before retrieval though, back when the Agent decides what's worth writing down at all. You can't pull clean signal out of a junk drawer. Get that right and retrieval mostly takes care of itself.
The cost angle makes selection even harder to argue with. You pay for every token whether the model reads it or not, so a bloated context bills you twice, once in dollars and once in the per-step errors all that noise creates.
Trimming what you load is a cost cut and a quality fix at once.
@CobusGreylingZA Every one of these is a new name for the same unsolved problem: deciding what the model should be looking at right now.
Prompt, context, harness, fleet, the label changes but the problem underneath doesn't. It's selection all the way down.
@DrJimFan Incredible setup!
The part I'd watch is day forty, when these Agents are still going and nothing stops them repeating experiments they already failed except whether the system remembered.
Autonomy at this scale is really a shared-memory problem.
@omarsar0 The orchestration layer is secretly a memory layer. Once multiple Agents coordinate, they need shared state, or each one forgets what the others learned and they drift apart. Agents that can't remember together just fail in sync.
@hwchase17 Stateful agent evals are overdue. The hard part is testing that it kept the right state, not just that it has state. The eval that predicts production is whether it still held the one fact that mattered 200 turns in and dropped the noise around it.
Four of these are one problem in different clothes.
The Agent has no reliable memory of its own work, and a bigger model fixes none of them.
Which of the five is costing you the most right now?
Five, it is most confident when it should stop.
The demo never shows the malformed input or the 3am moment it should have said I am not sure and waited.
An AI Agent that is 95% reliable on a single step drops to 36% across twenty steps.
The errors compound, they do not cancel.
Capability is not what breaks these in production. Four of the five real failures are the same problem.
10K Followers 3K FollowingExploring AI tools that make life easier & income smarter.
CPP: @Yapper_so & @Higgsfield
DM or Mail for Collab:
[email protected]
17 Followers 25 FollowingCMO ( Chief Meme Officer) @intraverse_game | I also help building @intraverse_game | Crypto Lover | Math lover | Music Producer
67K Followers 773 FollowingBreaking down the best AI tools & workflows | Helping AI startups reach the right audience | Product reviews • Launches • Tutorials
📩 DM for Collab. Mail 👇
311 Followers 228 FollowingThe Based Indian Community for Taiko ZK-EVM 🥁
,
It's not Official page from Taiko 〰️〰️〰️〰️〰️〰️
Explore Official Taiko :- https://t.co/JGkEeOSfr5
791K Followers 103 FollowingFirst based rollup in production. Home to AI Agents.
Proving Ground↓
https://t.co/7YhmXcpXr0
Submit a proposal ↓
https://t.co/26ceQJ22ov
222K Followers 3K FollowingFollow for posts about GitHub repos, DSPy, and agents
Subscribe for top posts
DM to share your AI project (Due to volume of DMs I'll prioritize subscribers)
10K Followers 1K FollowingFounder CEO #Wisekey and https://t.co/0emFmxYNzj & https://t.co/SeTENxXcKV , co-writer #transhumancode bestseller, former UN Expert #Cybersecurity WEF New Champion https://t.co/mKJLQm6jdF
10K Followers 2K FollowingGrowth & Global Banking @FlexSuperApp // Advisor @MoonPay // previously launched one of the first crypto cards (scaled to $500k in spend)
592K Followers 3K FollowingNVIDIA Director of Robotics & Distinguished Scientist. Co-Lead of GEAR lab. Solving Physical AGI, one motor at a time. Stanford Ph.D. OpenAI's 1st intern.
67K Followers 773 FollowingBreaking down the best AI tools & workflows | Helping AI startups reach the right audience | Product reviews • Launches • Tutorials
📩 DM for Collab. Mail 👇
30K Followers 7K Following🇨🇭 Not Forbes 30 Under 30.
Head of Developer Relations at @MetaMask.
Founding Venture Partner at @OuiCapital 🌍
Prev @digitalasset @angelHack
170K Followers 2K Followingtaking a break to self study physics. previously: early @coinbase, cofounder @scalarcapital and @bountycaster, dev ecosystem @farcaster_xyz
34K Followers 5K Followingi write about AI for Bloomberg @Technology. rachelmetz.11 on signal. she/her. opinions my own. [email protected] (tips yes, pitches nono).