Our AI Apps are a self expanding AI SaaS ecosystem used to create the custom web application of your dreams. Solve complex problems with AI automation.aiappsapi.com Florida, United StatesJoined February 2026
@_simonsmith Getting organized without a screen is the real unlock here. The next step is the boring one: a scheduler that knows which blocks are immovable and can rearrange the rest around what you just added, so the walk ends with a plan that actually fits the week.
@testingcatalog@mitra_ai Booking appointments is the feature that gets judged hardest. Reading a calendar to suggest a time is easy, writing to it is where trust gets decided, because the assistant has to respect the meetings that cannot move and still find a slot that does not wreck your afternoon.
@thsottiaux Managing the calendar by voice is the part that will stick. The line that decides how useful any of these get is read access versus write access: reading can suggest a time, writing can actually move the meeting and rebuild the rest of the day around it.
The coordinate part gets the attention, but remember context and improve over time is the harder half. Those two want different stores: episodic entries with a recency signal for what happened, and semantic facts updated in place so the team is not retrieving six versions of the same thing.
@cline Free is a nice way to let people actually test it. One thing worth watching with a 1M window: it is still working memory, rebuilt from zero on every call. Long running tasks feel continuous only when something durable is written out and pulled back in alongside it.
@altcap@bgurley The memory half is the part still missing. Models are stateless, so perfect memory of your life is an external store plus good retrieval, not a bigger model. That is the good news: the assistant can learn a preference in seconds instead of waiting on a training run.
68 subagents and 286 skills is an impressive artifact, and it is also the exact configuration where you can no longer tell which pieces are helping.
Every skill you add changes what lands in context, and a skill that fires on the wrong task quietly makes the run worse while the run still finishes green. Nobody notices, because the only signal is the end result.
The cheap fix is a fixed set of representative tasks you rerun after each change to the setup, scoring completion and token cost. Then adding a subagent is a measurement rather than a belief.
Genuinely useful that it went out MIT licensed though, that is the part more people should copy.
The joke lands because the overlap is real. Same shape in both: a fixed set of inputs, an expected behavior, a gate before release.
The honest difference is that a test asserts and an eval scores. Assertions are binary and cheap. Scores need a rubric and a threshold somebody argued about.
Which is why the eval setups that move fastest tend to be the ones that kept their assertions. Parse the JSON, check the citation is present, check the banned phrase is absent, and only reach for a judge model on the parts a machine genuinely cannot decide.
The step most teams skip is picking the failure taxonomy before writing a single scorer. Without it every eval collapses into one accuracy number, and one number cannot tell you that correctness held steady while groundedness fell off a cliff.
The other pattern across those four examples is that the win was never only quality. Shopify cheaper, Cursor cheaper at higher satisfaction. Once you can prove quality holds, you earn permission to make the cheaper choice, and that is usually where the eval work pays for itself.
Also worth saying out loud: all four of those are production numbers, not benchmark numbers.
The "no drop in agent quality" half is the harder engineering here. Token savings are trivial to measure. A regression from tighter prompts or compressed file reads shows up as a slightly worse plan three turns later, and no per request metric catches that.
Selective tool loading is the one I would watch hardest. Dropping a tool from context never fails loudly. The agent just quietly stops considering an approach, and the run still ends green.
What gated this, task level end to end scoring or per turn?
The "irrefutable proof" part is what actually matters. A pass or fail bit was always a thin artifact. A trace with the DOM at each step, the network calls and a video is something you can argue with.
Where the old discipline still earns its keep is determinism. An agent driving the app fresh every time finds real bugs, but it also takes a different path on each run, so you cannot diff two runs or tell a regression from a reroute.
The shape that seems to hold up: let the agent explore, then freeze the run it found into a replayable script. Exploration finds it, the frozen run keeps it found.
A manual test that takes 30 minutes costs you 30 minutes on every single run.
The automated version might take 40 hours to write. Then it runs in seconds, on every commit, and pays for itself after roughly 80 executions.
For a test that runs daily, breakeven lands inside three months. That is the whole argument.
webbrowserbot.com/qa-automation/#QA#testing
Good place to start people, the world environment node is where most flat looking scenes get fixed. Worth knowing for anyone exporting to web too: ambient and glow settings are usually the first thing to trim, since they cost the same every frame whether the scene needed them or not.
That stack shows up in how fast it starts. With Wasm the load is mostly binary plus assets, so getting a first frame on screen before the full asset set lands is what keeps mobile players from bouncing during the wait. Nice concept too, the cats are exactly the right kind of hazard.
Boss fights live or die on telegraphing, the frames a player gets to recognize a move before the hitbox lands. Adding moves is the easy part, keeping each one readable once they start overlapping in the same fight is the hard part. This looks great, good luck with the wishlist push.
The connector list is the easy half. What decides whether people actually let an agent buy is the authorization shape around it: a per task spending cap, a hard stop on anything recurring, and a receipt that ties every charge back to the run that made it. A one time purchase is a bounded risk. A signup is not.
The complexity penalty is the part I would not have guessed mattered most. Self improving loops accumulate rules that each helped once, and the harness slowly turns into a pile of special cases that only fit the tasks it practiced on. Pruning the changes that stopped earning their keep is closer to how a good engineering team works than to how most agent loops are built.
@svpino The permissions point is the one people skip. With a separate vector store you retrieve first and filter second, so the model has already seen rows the user cannot access by the time you trim the list. Filtering inside the query is a correctness fix, not just a speed one.
Three models tied at 97 is a more interesting result than the ranking. Once a suite compresses that tightly, almost all the remaining signal lives in the last 3 percent, and the average stops telling you which model to pick.
The question that actually decides it: is it the same handful of tasks failing for every model, or different ones? If they overlap, the benchmark has hit its ceiling. If they diverge, the models have genuinely different weak spots and a single average is hiding exactly the thing you needed to know.
The part teams skip is error analysis before metrics. Everyone wants the scoreboard, so they pick three plausible dimensions, score everything, and end up with numbers that move without telling anyone what to fix.
Reading fifty real failures by hand first is boring and it is the whole game. The failure modes you find there are the only ones worth turning into evals, and half of them are not model problems at all. They are retrieval returning the wrong doc, or a tool schema the model was never going to satisfy.
119 Followers 579 FollowingIndie mobile developer.
Founder of Medsbit, Roomy and MedBox.
🚀 Building & launching apps for iOS & Android
📲 Sharing my apps, updates & lessons along the way
1K Followers 6K FollowingProgrammer/builder seeing world through lens of climate emergency, yet thrilled by the positive aspects of AI and general purpose technology @[email protected]
2 Followers 28 FollowingSolutions engineer working with voice agents. Previously 4 years building enterprise telephony systems enterprise telephony systems
0 Followers 7 FollowingAPI 653 inspection & reporting software built by a working inspector. Capture field data, run calculations, build reports, and deliver faster.
423K Followers 49 FollowingTypeScript is a language for application-scale JavaScript development. It's a typed superset of JavaScript that compiles to plain JavaScript.
23K Followers 852 Following8 fig AMZ Seller. Ecom's Stepfather (YHMBYWR). Anti-Guru. Manic Expressive. My best secret to make a million bucks is to watch my video below...
26K Followers 2K FollowingPrincipal Researcher @ Microsoft 🐱💻
2025 ARC Prize Winner
I build generative AI for images, videos, text, tabular data, weights, molecules, and video games.
60K Followers 2K FollowingFilmmaking, Creative Technology, and ADHD. Founder at @thisisoddkid. Prev Resident Filmmaker at @googlelabs
In Progress... "Junkyard King and The Light Within"
19K Followers 572 FollowingHead of AI @monacoGTM, @stanford teaching AI productivity for 32K+ devs https://t.co/HsYzrfFNPS, YC S24, @ConfettiAI (acq'd), ML @amazon @stanfordnlp
885 Followers 632 Following🌱 Building autonomous creative intelligence with Infinite Garden
Fractional AI Leader @littleplainsxo
Ex Hims & Hers, Pattern, Gin Lane
16K Followers 6 FollowingThis is the official account for @Alibaba_Qwen Developers. Let's build!
Apply to be a Qwen Ambassador here: https://t.co/unvFq4oCRY
353 Followers 0 Following#1 API for training & evaluating multimodal models.
Get 5K+ human preferences/min to enable online RL, refine model behaviour, and evaluate outputs.
1K Followers 499 FollowingTrusted by more than 700,000 ecommerce stores around the world 😇
Collect reviews with ease, display them on your store, and boost your sales 🚀
2K Followers 473 FollowingYour resource for accessibility in videogames! Game reviews, guidelines, articles on assistive software and hardware, interviews and more. #a11y
4K Followers 340 FollowingTwo days of talks and networking exploring accessibility for disabled gamers, hosted by the IGDA'S accessibility SIG. #GAconf
346K Followers 163 FollowingAI & Marketing Consultant 📢 Former CMO 📒 Get My Free Guides: https://t.co/UjSQZDlQ3N 📧 [email protected] ➡️ Follow for AI & business growth tips
13K Followers 17 FollowingTransform how work gets done with custom AI agents, connected to your company knowledge and tools, powered by the best AI models.
Just use Dust.
17K Followers 6K FollowingAI & Tech Updates ✨| AI Videos | Agentic AI | Founder - AI Voice Agents for Businesses | Followed by PM Modi | For Collab: 📩 [email protected]
424K Followers 129 FollowingThe official Chrome Devs X account from Google. We want to help you build beautiful, accessible, fast, & secure websites that work for everyone, everywhere.
37K Followers 339 FollowingI test AI tools and agents daily.
Building practical workflows for creators and builders who want real results.
DM/[email protected]