This was an older post from me, @Keleesssss and few others at @datadoghq wrote earlier this year where we built multiple low level systems with opus 4.5 using a harness that involved formal verification (TLA+), deterministic simulation (DST) and Datadog observability working together in a closed loop. If anyone is interested in comparing notes, blog link in comment.
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached.
TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt.
There isn’t one definition of “good enough” for an AI application.
At @datadoghq's SF Summit, Han Lee of @Moodys put it simply: “It depends.”
Sometimes success is deterministic: did the test pass?
Sometimes it’s harder: did the customer actually get a good resolution?
Tony Hernandez of @JLL added another dimension: risk.
A lower-risk application with a human in the loop can tolerate a different standard. A highly autonomous system? “That’s a completely different ballgame.”
Start with: what happens if this gets it wrong?
Then set the bar accordingly.
AI agents can now investigate smarter — and spend less doing it.
Today at #DatadogSummit San Francisco, we announced the GA of Code Execution for the Datadog MCP Server: more accurate answers with 73% fewer input tokens and 40% fewer tool calls.
See how it works: bit.ly/4jjEsxA
The hard part of using AI to dismantle a monolith isn’t getting it to write the code. It’s proving the code is correct.
At @datadoghq's Summit SF, Sr Software Engineer Nick I. shared how a team of three is using agents to migrate routes out of Dogweb, our ~2,000-route Python monolith - recently crossing 100 routes migrated in a month.
His advice for agentic engineering workflows:
“How do I know that this code is correct without having to read it line by line?”
Start with the proof you’d need to approve the code, then work backwards to how an agent can gather that proof safely.
Happy to partner with @datadoghq - @datadogdevs for Boba & Build Night: Infra, ML, and App Dev in the Observability Age during #SFTechWeek!
Join us on October 8 for an evening of building, learning, and connecting with the SF tech community.
RSVP on this link: partiful.com/e/SiEiTLuTE847…
You can’t debug an AI agent in isolation. You have to debug the system around it.
At @datadoghq's Summit SF, @figma's Sr SWE Scott Gonyea explained that the agent’s execution record - prompts, model responses, tool calls and outputs - got the team pretty far when investigating failures.
But not far enough.
“When you’re investigating a user issue, you also want to know what’s happening in the browser session, what’s happening in the backend, in the requests and the logs.”
Because the failure might not be the model. It could be a browser issue, backend service, request, rate limit, schema change, or another dependency the agent relies on.
That’s why Figma built AI-Trace to bring that context together for engineers.
As AI products become more agentic, debugging them starts to look less like inspecting a conversation and more like investigating a distributed system.
“Now we’re moving to the next level where AI is more persistent, more independent, and doing more things.” - @embirico from @OpenAI
The next shift: agents that don’t just assist a user, but run as themselves against well-defined tasks the company can measure.
— @datadoghq Summit SF
One useful pattern for agent adoption:
Individual experimentation → repeated workload → managed agent → measurable outcome.
@OpenAI's @embirico pointed to testing, code review, incident response, and post-deploy validation as examples of work that can move from personal productivity into measurable automation.
Live from @datadoghq's SF Summit
What made Codex take off?
@OpenAI's @embirico broke it into 3 pieces:
- The model.
- The harness around the model.
- The user interface that sets the right expectations.
A useful framework for anyone building AI products.
@datadoghq SF Summit
“We’ve always been, at @OpenAI, basically value maxers, not token maxers.”
@embirico on why more AI usage doesn’t necessarily mean more value - in conversation with @datadoghq CPO Yanbing Li at SF Summit.
emerging from a 3 year (!) hiatus to share this. Code Execution is now GA in the @datadoghq MCP Server: 73% fewer input tokens, 40% fewer tool calls, and more accurate answers from every model we tested. run, don't walk! --> datadoghq.com/blog/datadog-c…
“Latest advances with our MCP is the ability to execute code. We give your agent a sandbox, with all the context Datadog has, to build the verification loop.” - Michael Whetten, Sr VP Product
@datadoghq
“With Agent Console, see all the AI that’s deployed and attach it to spend, to agents and AI infrastructure, because you have to guarantee your return on investment.
Not only track money spent, but tie it to work done.” - Michael Whetten, @datadoghq
Your logs already have the answers. Tap to Parse just makes them searchable — one click, no regex, no waiting.
Available now in Log Explorer and Log Pipelines, and upon request in Observability Pipelines. Read our blog to learn more: bit.ly/4yMKf39
“Building things has never been easier but our jobs, as builders, are not getting simpler.
Datadog was built to solve complexity and adapt to changing technology shifts.” - Adam Blitzer
@datadoghq
82 Followers 2K Followingbuilding to win 🚀
Co-founder and CTO of @kanawaiai
Former @google @cisco @forescout
@USC Marshall MBA
@UMBC Masters in Cybersecurity
927 Followers 6K FollowingBuilding @Cars24, Previously @Zomato. Building the future of mobility. Passionate about gadgets, cars, travel, space exploration & obviously good food. 🚀
349 Followers 3K FollowingPrivate equity real estate & AI. @columbia. 🇮🇪 in 🇺🇸. Interested in intersection of tech/AI and real estate investment industry. DMs open. @arsenal