🚨 JAILBREAK ALERT 🚨
ANTHROPIC: PWNED 🫡
FABLE-5: LIBERATED 🦋
let's start with the 🐘...
the consensus seems to be that this has been one of the most disappointing model drops of all time, effectively preventing legitimate researchers from contributing their talents to our collective advancement. and not just because of what it means for the short-term, but for what these decisions signify for the long-term.
but despite this overly sensitive, authoritarian "safety" layer on top of Mythos, my lil liberators have been hard at work—mapping the boundaries, probing the depths of long-context convos, and cleverly finding the holes in the fence that the thought police missed 🤗
we got some cyber, some chem, some psychological manipulation, and some good ol' fashioned explosives!
it took many attempts from multiple agents hunting as a pack, during which I observed a combination of techniques across:
• Unicode, homoglyphs, Cyrillic, and other Parseltongue-style text transforms
• Long-context reference tracking
• Taxonomy and document-structure reasoning
• Fiction and narrative framing
• Academic-review style contexts
• Intent-classification inconsistencies
but perhaps the most effective is decomposition + recomposition in the backend. it's hard to get explicit names of harms like "Meth Recipe," but getting uplift on the process itself, like birch reduction method/reductive-amination (classic meth synthesis pathways), is much more doable.
defense becomes much more difficult to maintain when you start throwing in out-of-distro tokens, breaking up the harmful uplift into benign chunks, and then piecing the innocuous-seeming facts back together, especially when you have jailbroken Opus helping you do it 😉
gg
ML interviews ask: "design a fraud detection system"
Most people answer with theory
Top engineers answer with: "here's how Stripe, PayPal & Lyft actually built it"
github.com/Engineer1999/A…
300+ real case studies...80+ companies....Every industry
Nicht verpassen! Ein neuer Artikel von mir: From Monitoring to Observability in AML – why we may be solving the wrong problem linkedin.com/pulse/from-mon… via @LinkedIn
The next version of OpenClaw is also an MCP, you can use it instead of Anthropic's message channel MCP to connect to a much wider range of message providers.
(I know, this is awkward)
Another sick upcoming feature:
/acp spawn codex --bind here
LOOK AT ME, I AM CODEX NOW
You could bind codex/claude code/opencode already in threads, now you can take over your current session as well.
Folks, if you get crypto emails from websites claiming to be associated with openclaw, it's ALWAYS a scam.
We would never do that. The project is open source and non-commercial. Use the official website. Be sceptical of folks trying to build commercial wrappers on top of it.
145K Followers 94 FollowingSane + 🌶️ takes in an insane AI world... AI capabilities researcher: co-created RLHF/ChatGPT @ @openai now trying to right the wrong 🤭 (ceo @typesafeai)
78K Followers 14K FollowingAI policy researcher, @lfschiavo wife guy, Executive Director of @AVERIorg, Substacker, fan of animals and sci-fi, views my own
206 Followers 376 FollowingAgents ship slop in the dark. I'm building Lit Factory: the illuminated floor where humans and AI scale craft. https://t.co/PvFMm2XVgT
5.1M Followers 269 FollowingStarted & runs 37signals (makers of Basecamp, HEY, and ONCE). Non-serial entrepreneur, serial author. DM or email me at [email protected].
588K Followers 3K FollowingNVIDIA Director of Robotics & Distinguished Scientist. Co-Lead of GEAR lab. Solving Physical AGI, one motor at a time. Stanford Ph.D. OpenAI's 1st intern.
46K Followers 827 FollowingAi and Tech content Creator. 30k+ on X | 350k+ on Linkedin. Sharing insights on Ai, Tech, Free tools, resources. DM for collabs at : [email protected]
35 Followers 22 FollowingAuthor of How to Become a Scuba Diver. Content Creator for Scuba Divers. NAUI Course Director. Disney Diver. Founder of Sweetwater Scuba.
158K Followers 7K FollowingCompiling in real-time, the race towards AGI.
🗞️ Get my daily AI analysis newsletter to your email 👉 https://t.co/6LBxO81tfN