the design of evaluatorbench.com grew out of a summer project, epistemedia.org (which i still need to write about), and runs on @activegraphai
the pattern across all three projects is that it treats the journey as the primary substrate, and the destination as a projection of it
the independence of AI evaluators is going to matter more... so i made it a "benchmark":
evaluatorbench.com (research preview)
source-linked directory of third-party evaluators of frontier AI. each one has an independence score you can take apart: change the weights,
found a new @ActiveGraphAI citing in a new paper on graph engineering for llm agents ☺️
arxiv.org/pdf/2608.21156
"Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence"
🚨 STOP RE-RUNNING YOUR WHOLE AGENT TO TEST ONE CHANGE.
someone just open-sourced the fix. free
activegraph. 547 stars. one install.
install it - and your agent gets:
- fork a finished run at any step
- the shared part replays from cache - zero new model calls
- strict replay re-fires every step and fails on the first divergence
- an append-only log: what changed, when, and why
- policies hold an action until a human approves it
setup, 30 seconds:
pip install activegraph
activegraph quickstart
runs on recorded examples. no keys, no config, same output every time
save it before your next long run dies with no log to replay ↓
[new @activegraphai blog post] How Synthetic Players used ActiveGraph to verify 4,919 runs without new model calls
activegraph.ai/blog/synthetic…
For the increasing number of agentic academic papers, reproducibility, etc become valuable in strengthening your argument
This blog post walks through how we used ActiveGraph for this.
there's an increasing amount of research replacing LLMs as human subjects, so I was curious how well this actually works
to do this, i had LLMs play well known games, and compared how they played, compared to real humans
arxiv: arxiv.org/abs/2608.00979
site:
there is a missing link between how humans and agents collaborate today. the atomic unit of focus is a "task". but there's no way to merge tasks together in a way that gives context over the entire "workstream".
@yoheinakajima's @ActiveGraphAI has the right model of an append only event log, but theres need for a different way to do projections.
ive been thinking about how to take an event stream and "fold" it into something that enables better human-agent collaboration.
human - agent collaboration is still pretty sloppy.
- chats and sessions are fine, but they tend to be extremely unorganized.
- artifacts aren't created in a reasonable location, they're not visible
- the ui and framework used is owned by a model company (eg claude cowork)
-
Nine iterations of BabyAGI in three years, and yet the bit that @yoheinakajima kept coming back to was graphs. @aiDotEngineer published where that landed, "Active Graph Agent Runtime (BabyAGI 4)", on YouTube. The legend Yohei is Managing Partner at Untapped Capital.
The talk is a working argument for building an agent around an immutable event log instead of around the LLM, with code, reference agents, and experiment results behind it.
- The log is the agent. One immutable typed event log holds what the agent did and every change to the agent itself, and it projects the graph that is the agent's state.
- Behaviors replace the control loop. They react to graph changes and emit events. LLMs never talk to each other, only to shared state, so replay, rollback, and forking come natively.
- Policies decide what can change. Adding a research source is cheap. Editing a prompt can require a human. A new fact can require that nothing contradicts it.
- Views are context management as a graph query, Handing a behavior the subset of the graph it should see.
- Packs, not skills. Memory, identity, tools, secrets, chat: each bundles object types and behaviors, so you can swap one memory pack for another.
- A runtime, not a harness. He rebuilds ReAct on top of it. On goal created, add a thought. On thought created, run reason.
- The log doubles as memory. On LongMemEval, no fact or entity extraction, just embed the query and pull the messages around the hits. When his API key ran out at question 350, the run picked up at 353 instead of starting over.
- Self-modification with gates. Regimes classifies the failure, lets the agent edit only the matching part of itself, then requires a static check, a sandbox check, and a measured rerun before a patch is accepted. Loops of 8 to 13 accepted 4 or 5, with modest but statistically significant gains.
- It remembers what failed. Around 80 tuning passes on a deterministic Pokemon trading card agent for a Kaggle competition, 20 to 30 accepted, and everything that didn't work stayed on the record.
- Old architecture, new workers. Blackboard and Kafka have decades of writing behind them while LLM agents have three years, which is his hypothesis for why coding agents write this style well.
I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!
Yohei Nakajima, the creator of BabyAGI, just showed why he stopped building agents around the LLM.
His AIE talk on ActiveGraph, in 5 timestamps:
1:55 – build around the log, not the model
3:24 – behaviors and policies gate what the agent can change
8:12 – his api key died at question 350, the run resumed itself at 353
11:11 – a loop that forks the agent and keeps a patch only if accuracy rises
15:49 – why long-running agents need an experiential world model
The frame: make an immutable event log the agent, and replays, rollbacks and forks come for free.
Different way to think about agents. 17 min well spent.
Anthropic's Lance Martin just walked through how they build agents that run for hours with no human in the loop.
His talk on async, long-horizon agents in 5 timestamps:
3:19 – split the brain from the hands
6:20 – build agent vs verifier agent in a loop
8:18 – 20 iterations to
having an event stream like activegraph is the base of how we move to next level and that will probably be composition, which can let you achieve better results with smaller models.
- get the core event stream
- wrap things in a nice monad
- now you compose the models:
> main runner's stream
> add on a validator agent that checks specific user interaction and changes/obfuscates data
> parallelized tasks that run multiple agents at same time
> classifiers, mappers, tinkerers, translators
and some can be SOTA models, some can be super fast models, some can be local mini models.
and this way you can build highly complex and observable pipelines, "swarms" or "brains" for whatever you need. much more solid, reproducible and observable than your generic imperative orchestration.
🆕 ActiveGraph: The Log is the Agent
my talk from AI Engineer is live!!! 😆
youtube.com/watch?v=khVX_B…
it's about @ActiveGraphAI, an event-sourced graph runtime for building durable long-running agents:
- what it is
- code examples
- experiments
- benefits & surprises
please
I've long felt AI harnesses were trapped by our obsession with code so @yoheinakajima's ActiveGraph video hit me hard. Reminded me of time I obsessing over Linda space & case-based reasoning. Watch the video. It's very different.
youtube.com/watch?v=khVX_B…
🆕 ActiveGraph: The Log is the Agent
my talk from AI Engineer is live!!! 😆
youtube.com/watch?v=khVX_B…
it's about @activegraphai, an event-sourced graph runtime for building durable long-running agents:
- what it is
- code examples
- experiments
- benefits & surprises
please check it out, share with your friends who like to tinker at the frontier, and let me know if you have any Qs!
your hippocampus fast-captures events (logs) that are computed on to hold a current state, and through selective replays, pulls some of this slowly into the cortex (model), which then feeds new priors back into the computations
continued working on @activegraphai reference packs last weekend, resulted in needing to harden the runtime:
it already could...
- keep a complete history of everything an agent did
- replay that history
- fork into alternate timelines to try different ideas
it can now...
- realize it's missing a capability
- write new code to add that capability
- test it in an isolated copy of itself
- ask a human to approve it
- safely adopt it
- continue running with the new capability
activegraph.ai/blog/activegra…
the other week i read active graph papers from @yoheinakajima and got inspired.
enough to rewire brigade, graphtrail, and miseledger around one habit: state is a projection of a log.
then i got a chance to read this yesterday and the same receipts turned out to be mineable. brigade now finds command sequences operators keep repeating, proposes runbooks from them, and pins each step's binary by sha256.
the whole loop: brigade.tools/blog/activegra…
this is a great approach, seeing this more
@FlyMy_AI also does this when you build an agent via their api, they'll build a deterministic reusable workflow, except for where you need models
here it is: brigade.tools/blog/activegra…
ended up as the three that close a loop: graphtrail diffs the code graph, the diff rides into brigade's run receipts, exports to miseledger as content-addressed evidence, and comes back as context for the next run.
one wrinkle: a single event log wasn't enough for the promotion ratchet, it needed the decision receipts as their own transition log.
207 Followers 314 FollowingHyper growth companies, working with great people, driving opportunity, leading by example, instinct for change and passion for life.
5K Followers 3K FollowingBuilding #ElectricClojure and https://t.co/JkYr6J4dpu. I believe in excellence, and I believe that many others do too. Baháʼí. https://t.co/aBmRZmuoCs
130 Followers 2K Followinghusband and father of two girls. CPA. currently Controller at PDA, Inc (healthcare consulting firm) and Controller at Shoal Harbor Capital (PE firm).
788 Followers 2K FollowingI'm virtual, see. Yeah, see, yeah. Software, dad jokes, public speaking, and other such business. Lead Software Engineer @focused_dot_io
2K Followers 2K Followingבעיקר כדורגל וקצת טכנולוגיה. הפועל פתח תקווה, ואז אתלטיקו מדריד ו-ארסנל. כתבתי ודיברתי ב- @HazavitIsrael על כדורגל. בונה דברים ב- @salto_io
2K Followers 1K FollowingLeading engineering at @LeptonSoftware..Trying to build digital twins on the web..and making it easy to build your own tech stacks with Vinxi
526K Followers 6K FollowingCo-founder & CEO of Discovery Loop. Former Chief Scientist, Google. Helped build many Google products, TPUs, Gemini, TensorFlow, MapReduce, Bigtable, ...
84 Followers 278 FollowingBuilding agents & autonomous software infrastructure. Lead AI @GetStimulus | Databricks & scalable AI systems expert
Life · Arcan · Agentic Control Kernel
298K Followers 857 FollowingCooking fun AI systems & products @databricks. Prev: co-founder & CTO @ Hyperbolic, OctoAI (acquired by @nvidia) Apache TVM, PhD @ University of Washington.
1.9M Followers 1K FollowingCo-Founder of Coursera; Stanford CS adjunct faculty. Former head of Baidu AI Group/Google Brain. #ai #machinelearning, #deeplearning #MOOCs
1.9M Followers 179 FollowingNobel Laureate. Co-Founder & Chair @GoogleDeepMind; Chief Scientist of Alphabet. Founder & CEO @IsomorphicLabs. Building the future...