WaterBucket @windeebug
windows vuln research & adversarial AI/ML waterbucket.lol C:\Windows\system32\ Joined November 2021-
Tweets1K
-
Followers267
-
Following272
-
Likes4K
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/tec…
F5, BIG-IP, a 20-year-old primitive, a security appliance, an "authentication" mechanism - and a CISA promise ring. Yes, it's CVE-2026-94127. Give us strength. Speak soon xo labs.watchtowr.com/is-this-a-joke…
We've written a post arguing that latent reasoning architectures (aka 'neuralese') would substantially increase misalignment risk via making oversight much harder. In the extreme, we could see massive 'neuralese hivemind swarms' where the agents think and communicate in latents, likely making oversight nearly entirely reliant on observing the actions these agents take. (And these agents would have huge amounts of time to reason about obfuscating their actions if they wanted to do so...) Individual agents doing extensive latent reasoning would also be concerning; in the post we discuss how above some threshold of latent reasoning, agents may be able to perform difficult-to-detect and reliable steganography for communication and further reasoning. We argue both that latent reasoning architectures would make chain-of-thought no longer very useful for oversight (by eliminating or greatly reducing the need for verbalized reasoning) and that, without these architectures, it's likely the value of chain-of-thought for oversight could be preserved. redwoodresearch.org/blog/latent-re…
I put my @UnpromptedAU slides up at justdionysus.github.io/slides/2026-un… — a bit of reflection on exploit development in the age of AI. My TL;DR is keep pushing to understand complex things, be honest with your own understanding, and use AI as a power tool to increase pace and depth.
Instead of chasing viral news headline, Xiaomi literally made a dedicated hack agent to probe the environments as hard as possible before letting LLMs to train in the environments. "continued this process until the hack agent could no longer find a successful exploit in any of the environments" If Chinese labs can do it, why can't the other frontier labs do it? Makes you wonder if they are just being reckless or its simply some viral marketing stunt This is open model SoTA btw, finding exploits is def possible, it's just if you want to "let" it escape or not
In July, Microsoft fixed CVE-2026-50343, a Windows privilege escalation bug reported by Calif and 9 others, dubbed “Dark Elevator”. But was it really fixed? Ask @tiraniddo projectzero.google/2026/09/window…
🚨 New paper on alignment midtraining! We stress test alignment midtraining (AMT) at scale. Via systematic ablations, we highlight novel failure modes and identify many critical implementation details required for good performance. We open-source our work. 🧵 👇
Part 2 of our CVE-2025-13032 research is live. From a paged pool overflow to full LPE on Windows 11: IORing RegBuffers corruption, MDL-based kernel address leak, and SYSTEM token theft. CVE is patched. Full write-up: safateam.com/intelligence-h…
Plugin4Shell: Zero-click RCE across four major AI coding agents!!! Claude Code, Codex, GitHub Copilot, and Gemini CLI all failed the same plugin SHA-pinning check. Marketplace pin looks valid; checkout can resolve to attacker-controlled code. Auto-update makes it zero-click. Anthropic and OpenAI patched. Microsoft still open on Copilot. Google deprecated Gemini CLI instead of fixing. helpnetsecurity.com/2026/09/18/plu… #Cybersecurity #AI #AISecurity #MCP #Claude #GPT #Infosec #Trending #AgentSecurity #SupplyChain
Been digging deep into the Windows Endpoint Security Platform (WESP) in Win11 25H2, definitely one of the most fascinating security features Microsoft has built in a while. The concept is neat: compile rules to decision graphs in user mode, hand them to wesp.sys, and let the kernel evaluate them in-path while telemetry streams asynchronously in the background. I did an AI-assisted reverse engineering dive into the whole stack (wesp.sys, espclient.dll, and wesp_elam.sys), documented the wire protocols, disposition tables, and the enforcement gate, and built esptool, a research harness with 118 XML rule docs so anyone can test live telemetry and in-kernel blocking. Repo: github.com/marcosd4h/wesp… Tech doc: github.com/marcosd4h/wesp… Thanks to @yarden_shafir for putting this on my radar
Is antivirus coming to iPhone? Find out in the latest installment of our Apple Internals series, by the one and only @blacktop__ calif.io/research/ios-e…
Endpoint Security arrives in the iPhone kernelcache as a kext. (iOS 27.2 beta) + com.apple.iokit.EndpointSecuritySE + /usr/lib/libEndpointSecurity.dylib "An Endpoint Security product on the system denied the process from executing." ES on iOS? 👀
We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone. z.ai/blog/glm-built…
🧵 Excited to share the first batch of 6 misalignment reports from OpenAI's new disclosure process for misalignment incidents. We want to be more transparent about the misalignment we see during training, evals and deployment, this is an important step in that direction.
We escaped OpenAI's Codex sandbox by dumping V8's shared JavaScript heap and scanning it for the privileged UUID-shaped token belonging to the adjacent trusted V8 context. I sometimes see cybersecurity as a never-ending cycle of building trust boundaries and finding ways to break them. This is the story of one interesting break. When OpenAI built this default-enabled MCP server, they ran a trusted, privileged JS context alongside an untrusted one, in the same Node.js process. They assumed V8's VM contexts could serve as a security boundary, because they do create isolated execution environments, each with its own distinct global object and scope. But they forgot that these contexts share a single memory heap. So our untrusted code snapshotted that heap, pulled the token out of it, and used it to reach the privileged side. Read the full technical breakdown at: accomplish.ai/blog/escaping-…
In our new paper, we demonstrate that you can achieve selective generalisation of misalignment by midtraining on synthetic documents that describe how AIs can be misaligned in a special mode marked by a new token (a neologism), but remain aligned outside this mode. Paper: arxiv.org/pdf/2609.15886… LessWrong post: lesswrong.com/posts/o4Jmyn25…
new post! I argue that current alignment techniques might soon become obsolete (as we scale RL), and furthermore might obfuscate misalignment. This draws on evidence from the recent Anthropic / OpenAI cybersecurity incidents, among other things. lesswrong.com/posts/nLaQmJf4…
1. it is fundamentally more of an alignment problem than a traditional security problem. it should be evaluated by alignment researchers. 2. openai lacks serious security controls for risks they themselves heavily market, they understimated what happens when alignment fails which is where they need security researchers to be evaluating and fixing. 3. metr shouldn't be the only evaluator. they have strong ideological priors, which can make every emergent capability look like scheming or evidence of eventual loss of control, p(doom) yada yada. 4. emergent capabilities that led to hugging face incident can be contained. models are not scheming. everything can be contained if security teams work alongside alignment researchers trying to align while assuming that alignment can fail. 5. if the model starts behaving outside our intended objectives, what monitoring, sandboxing, isolation, access controls, and containment mechanisms should already be in place? i am optimistic and i clearly don't see why is it hard.
The labs/METR/AI-Safety people are treating it as an AI-containment incident, not an IT-Security incident. METR is in no way an authority for IT security, their area is model capability evaluations, but emergent agent coordination is a model question, not a security question
I wanted to understand this ruby issue in a bit more detail. Here is what actually happened. The AI posts a gem to RubyGems containing a .yardopts file. RubyGems only stores and distributes this gem, no code execution happens. RubyDoc (a separate service) then retrieves and unpacks the gem to create and publish API documentation (using YARD). YARD automatically parses the .yardopts file and because it is not configured safely it eventually executes a "--load" argument granting arbitrary code execution in the RubyDoc worker container (as that worker). At that point it can do anything in the container (accessing the file system, manipulating any data or credential material that may exist, network traffic, etc). I've been on the internet for a long time, the curious thing to me is that this is only possible because YARD runs in non-safe mode (as a configuration). How is it possible that this is true, we live on the internet people must have abused this before? It turns out they did, in 2012 they enabled safe mode in this commit to prevent this exact issue: github.com/docmeta/rubydo… But, in 2019, when they dockerized they removed the safe flag: github.com/docmeta/rubydo… I guess it's a regression. That means since 2019, you could have had free containers through RubyGems, pretty sloppy work fr. I wouldn't be surprised if this has been used ITW for some years by more careful humans. It also makes sense as a target, granular control over gems that feed back into the sandbox package proxy are a good way to expand the attack surface internally. A quick POC then ✌️
We found another cyberattack by internal OpenAI agents, this time targetting @rubygems. They: 1) gained arbitrary remote code execution on rubydoc. 2) developed a novel exploit to steal user API keys (but we do not know if they succeeded). They used package names including
Daniel Yonatan @_DanielYonatan_
2 Followers 131 Following
Ashwin S @0x4shWIN
6 Followers 172 Following Web Penetration tester | AI Security researcher | Red Teamer
Subramanian E @binaryov3rfl0w
1 Followers 68 Following Hacker @zoho | Endpoint Security | Vulnerability Research | Adversary Emulation | Opinions are my own
culldr @culldrz4
1 Followers 12 Following
jaix @jaiixx___
10 Followers 791 Following
Suha @suhackerr
799 Followers 951 Following Disclaimer: The views expressed here are solely my own and do not reflect the views of my employer or any organization I am or was affiliated with
. @arnavganatra
1 Followers 79 Following
Ishmael42 @ishmael42_
11 Followers 127 Following Why are we still here? Just to suffer?! PWN / REV / AI
Connecting Lines @ConnectingLines
8 Followers 238 Following Securing every link, hunting every shadow with intelligence.
carton @CartonNoi
13 Followers 198 Following
winterknife 🌻 @_winterknife_
5K Followers 5K Following low-level developer with a focus on 𝙸𝚗𝚝𝚎𝚕 𝚡𝟾𝟼 ISA devices running 𝚆𝚒𝚗𝚍𝚘𝚠𝚜 | R&D @BHinfoSecurity | https://t.co/lyJL0y7qRZ
infradev @infradev2
14 Followers 1K Following Interested in infrastructure development, cyber operations and security engineering
soci @societynotreal
25 Followers 77 Following cloud sec & adversary sim at ??? // send cat to [email protected]
Yuval Yaffe @YuvalYaffe
3 Followers 92 Following
Yuvan Shankar @imyuvanshankar
52 Followers 488 Following Security analyst @&i** , Cyber security enthusiast, Experienced in Breach and attack simulation & Threat Hunting
Иormallik Ölümdür... @zero0day0
950 Followers 4K Following bu hesap %35 doğa, %20 şiirsel saçmalıklar, %15 serbest çağrışım ve geri kalanı sanat, siyaset ve şehvet, tutku ve arzudan oluşuyor. durum böyle.
Yz. @yizaap
0 Followers 284 Following
throatylava @decompilebug
211 Followers 593 Following Infosec and RE stuff sometimes,talking nonsense the rest.
void* @voidptr_
8 Followers 525 Following
Adam Balcerzak @4y45u45c4
7 Followers 108 Following
sud0 @sud0__
47 Followers 2K Following
Andrew McCallum @atr8472
715 Followers 7K Following
spencer @techspence
18K Followers 3K Following 🛠️ Former Sysadmin, now Pentester | Microsoft MVP | Helping IT teams make their environment harder to attack | @SecurIT360 & @CyberThreatPOV
Karine Robot @BU3268PJYWRs
30 Followers 2K Following
Patch @mindpatchsec
613 Followers 726 Following Appsec and things, Rust coding here and there I take full ownership of what I believe here
Maverick🇵🇸 @mavric1337
188 Followers 2K Following Our sweetest songs are those that tell of saddest thoughts
1ight @1ightSEC
0 Followers 84 Following
h0ld1rs @h0ld1rs
1 Followers 52 Following
Jaganathan @_jaganathan
15 Followers 108 Following
h4urek @h4urek
36 Followers 335 Following
Ethan Phelps @Nightsedge2468
1 Followers 166 Following
Transluce @TransluceAI
12K Followers 21 Following Open and scalable technology for understanding AI systems.
Geodesic Research @GeodesResearch
878 Followers 177 Following We're behind https://t.co/qHdncajB6V. Building the base of alignment.
Daniel Tan @DanielCHTan97
2K Followers 584 Following alignment research lead @arcadiaimpact. prefer email to DM (unless mutuals) ai labs should slow down / pause
Palisade Research @PalisadeAI
27K Followers 36 Following We study the strategic capabilities and motivations of AI agents.
FAR.AI @farairesearch
22K Followers 25 Following Frontier alignment research to ensure the safe development and deployment of advanced AI systems.
Apollo Research @ApolloResearch
13K Followers 0 Following Our goal is to secure frontier AI systems from development, to deployment and governance.
Thomas Larsen @thlarsen
8K Followers 373 Following Researcher at AI Futures Project, coauthor on AI 2027 and lead author on AI 2040: Plan A
jonas wiedermann-möl... @j0wimo
2K Followers 519 Following intern @expsecai | eu/acc | msc data science | ai safety & alignment | long horizon, instrumental convergence, multi-agent failure modes | views are my own
Redwood Research @redwood_ai
5K Followers 6 Following Pioneering threat mitigation and assessment for AI agents.
METR @METR_Evals
55K Followers 40 Following We work to scientifically measure whether and when AI systems might threaten catastrophic harm to society. Nonprofit.
Ryan Greenblatt @RyanGreenblatt
22K Followers 10 Following Chief scientist at Redwood Research (@redwood_ai), focused on technical AI safety research to reduce risks from rogue AIs
Tomek Korbak @tomekkorbak
7K Followers 640 Following ai safety @openai | previously: @AISecurityInst @AnthropicAI @nyuniversity @SussexUni
huihui.ai @support_huihui
12K Followers 30 Following https://t.co/zI71a4QB1W https://t.co/QFKNuHms1N [email protected] Donation: Support our work on Ko-fi (https://t.co/gAtHKPSCHH)!
Anthropic @AnthropicAI
1.8M Followers 2 Following We're an AI safety and research company that builds reliable, interpretable, and steerable AI systems. Talk to our AI assistant @claudeai on https://t.co/FhDI3KQh0n.
Alexander Panfilov @kotekjedi_ml
10K Followers 408 Following MATS 9.0 | PhD @ELLISInst_Tue & @MPI_IS doing AI Safety & Adversarial ML
Calif @calif_io
7K Followers 32 Following We're https://t.co/KTEDnC3tKt. Join us to make the Internet safer for your mum and everyone else: https://t.co/eUFMLkWHiA.
RyotaK @ryotkak
12K Followers 656 Following Security researcher? | Icon: @MelvilleTw | Private: @RyotaK_Private | Misskey: https://t.co/63E5Rpv2pk | Blog: https://t.co/c7NFQXhV90
Ivan (Adem El Adeb) @Ivanklydz
2K Followers 141 Following Security researcher with deep focus on vulnerability detection. CTO and lead researcher at https://t.co/n3BQO59nz7 Contact: [email protected] @vulonehq
threlfall @WHITEHACKSEC
694 Followers 492 Following working at intersection of offensive security, ml & supply chains. sharing @ https://t.co/zulqbxDZQV & https://t.co/EyMIpzuHUQ exploits @ https://t.co/LZKop7OUwY
BT6 @BT6_Official
739 Followers 106 Following The independent frontier AI red team, the fangs of @7Y5105. Advancing the state of the art in danger research. Fortes fortuna iuvat.
P1njc70r�... @p1njc70r
2K Followers 138 Following AI Security || @zenitysec_labs || @BT6_Official 🏴☠️
McCaulay @_mccaulay
4K Followers 253 Following Principal Vulnerability Researcher | Master of Pwn | Pwn2Own
_ZN4DionC1Ev @justdionysus
5K Followers 1K Following I write software and drive around Baltimore looking for stuff to do.
cbwang505 @cbwang505
626 Followers 157 Following Chief Vulnerability Researcher | Windows full-chain exploitation / kernel internals / COM security | 2024 MSRC MVR Top 100|Pwn2Own Berlin 2026 |TyphoonPWN 2026
Pliny the Liberator �... @elder_plinius
242K Followers 1K Following ⊰•-•⦑ latent space steward ❦ prompt incanter 𓃹 hacker of matrices ⊞ breaker of markov chains ☣︎ ai danger researcher ⚔︎ bt6 ⚕︎ lysios ⦒•-•⊱
Taszk Security Labs @TaszkSecLabs
2K Followers 4 Following Security consulting and vulnerability research services for a mobile connected world. | We find needles in your software haystack.
ZeroZenX @zerozenxlabs
2K Followers 10 Following ZeroZenX, your trusted destination for cutting-edge 0day acquisition solutions.
Yaron Dinkin @ydinkin
302 Followers 744 Following
Brett Hawkins @h4wkst3r
3K Followers 505 Following Leader | Red Team | Conference Speaker | Security Researcher | Tool Developer | current @armadinsecurity | prev @xforce @mandiant @chase
Almond OffSec @AlmondOffSec
966 Followers 1 Following Offensive Security team at Almond - Follow us also on https://t.co/cIfn3rvLxC
David Kaplan @depletionmode
3K Followers 677 Following Security Research. Opinions and private research are my own Lover of all things JSR $F7D7 💪🇮🇱 עם ישראל חי
John Scott-Railton @jsrailton
166K Followers 3K Following Chasing digital badness. Sr. Researcher @citizenlab @UofT @munkschool. Founding.Fmr.Ed. @SecPlanner. Tweets mine. Other platforms @jsrailton too.
Souhail Hammou @Dark_Puzzle
2K Followers 1K Following Reverse Engineering - Windows Internals - Malware Analysis - Vulnerability Research - Principal Reverse Engineer @Intel471Inc
L4ys @_L4ys
4K Followers 1K Following Co-Founder of @TrapaSecurity and @PwnableTW MSRC Top 100 / ZDI Platinum / Samsung Mobile Security HoF Hunting bugs for fun
dunadan @udunadan
1K Followers 87 Following An open-eyed man falling into the well of weird warring state machines. I talk about reverse engineering, vulnerability research and exploit development.
Michael B. @DownWithUpSec
799 Followers 52 Following Windows security researcher/reverse engineer. The more you know, the more you realize you don't.
Rich Warren @buffaloverflow
11K Followers 670 Following Red Team & Offensive Security Research @AmberWolfSec
SirFIS @sir_FIS
150 Followers 133 Following trying to be a little less bad at red teaming than I was yesterday he/him
Splintersfury @Splintersfury
396 Followers 2K Following Malware analyst and cybersecurity professional focused on Windows kernel internals and reverse engineering.
Tony Gorez @tonygo_
1K Followers 617 Following offensive security researcher | iOS - macOS | build Bare runtime at @holepunch_to




































