UPDATE: Kev-0.6B, 4B, and 8B are now available. Kev is a family of small open source Jev-like decision models you can train and run yourself.
This new family is based on Qwen3 using the same LoRA + small pointer head technique as before, but scaled up.
Out of domain, on data
The funniest part about local AI:
you start because you want privacy
A few weeks later you're comparing memory bandwidth, quantization formats and $5,000-$20,000 possible hardware purchases at 2am
Jev launched Monday as a closed API. By Thursday we already have an open weight alternative we can run locally.
This is fun, looking forward to test it!
huggingface.co/convaiinnovati…
Absolutely incredible work by @mdowd and @UnpromptedAU putting together an amazing collection of people and talks. Favorites were @chompie1337 and @justdionysus – great speakers with deep expertise in exploit dev. Hope to come back next year for another round!
Umbriel's Caleb Gross (@noperator) has an article in the latest Phrack 73 issue (@phrack) "Word Machines for Weird Machines" - check it out when you get the chance!
Earlier this summer, I bought a 128 GB Framework Desktop just so I would have access to a large continuous block of unified memory to run frontier-adjacent models at home. I've been really inspired by @antirez's DS4 project which aggressively targets specific model+hardware combos. I recently became aware of @ilintar's project doing this specifically for Strix Halo and I'm really excited about it.
@antirez I've been running DeepSeek V4 Flash 0731 through DS4 engine on the Framework. With a Q2-class quant, I was getting roughly ~200 tok/s prefill and ~15 tok/s decode. I did not feel that this configuration was fast enough to work well in an ongoing Hermes loop. Qwen3.8-Flash-Next is
I've been running DeepSeek V4 Flash 0731 through DS4 engine on the Framework. With a Q2-class quant, I was getting roughly ~200 tok/s prefill and ~15 tok/s decode. I did not feel that this configuration was fast enough to work well in an ongoing Hermes loop. Qwen3.8-Flash-Next is a much smaller model while still roughly on the same capability level of DeepSeek V4 Flash 0731 (even better on some tasks), and I'm running it at a Q4-class quant. Wilkin has been aggressively optimizing llama.cpp/ROCm specifically for Qwen3.8-Flash-Next on Strix Halo. On my Framework I'm seeing roughly ~1,300 tok/s prefill and ~30 tok/s sustained decode at around 131K context.
"If you wanted to win the 2026 Nobel prize in physics, you have to be a physicist: not a musician who dabbles in physics, or a politician who has a physics hobby in your spare time. You have to be fully immersed in the world of physics. AI research is not like this. We are very
@chrisrohlf Big 4 academic computer security conferences are dealing with this already. Every submission cycle mentions "the significant increase in the number of submissions."
Excited to announce that I recently joined @Umbriel_AI with @mdowd and @dyn___ :) Humbled and grateful to get to work with both of these talented hackers.
253K Followers 1K FollowingCofounder @hackinghub_io | Advisor @CaidoIO. I hack companies and make content about it. #NahamCon organizer. ex @hacker0x01🇮🇷
34K Followers 1K Following意志 / mobile research @ ▓▓▓▓▓ / Team 501 / ex IBM Capability Lead & FireEye TORE / I rewrite pointers and read memory / AI Psychoanalyst / BHUSA Review Board
322 Followers 386 FollowingRE/Vuln dev who now does some business things. Father of nine. Moving to using this strictly for news, DMs, and likes...trust me, I'm here
795 Followers 949 FollowingDisclaimer: The views expressed here are solely my own and do not reflect the views of my employer or any organization I am or was affiliated with
114 Followers 1K FollowingI do this not because it is easy, but because I thought it would be easy... Yeah man like what even are computers and why they need to be secure ( ‘• ω • `)?