Welcome back. Grok Bot is putting rival models to work – a surprising move for a company directly in competition with other AI labs. Meanwhile, Anthropic rolled out its smallest and cheapest model yet, and Google's latest Playground lets you build playable video games straight from a prompt.
Also: Microsoft revamps its software for agentic AI, a founder’s wiped C drive offers a costly Claude Code lesson, and Apple’s new smart home launches have devs hooked.
Today’s Brief
The eval playbook for anything you ship on an LLM
An app to organize your desktop screenshots
Build an AI software factory with Claude Code (tutorial)
When to load MCP tools (cookbook)

TODAY IN PROGRAMMING
SpaceXAI's Grok Bot borrows rival models for tough tasks: The Musk-led lab just turned Grok Bot into a model router that can route tasks to outside models such as Anthropic's Claude Opus 5.5 for reasoning, Midjourney for images, and Suno for music. Simple questions go to small, fast models, while trickier requests get the heavyweights. Grok Bot can also now search and monitor X without a connector, with early users reporting no API costs.
Anthropic ships its cheapest and fastest small model yet: The AI lab just dropped Claude Haiku 5.5, the first Haiku model that lets you dial its effort up or down. Adaptive thinking is on by default at medium effort. Built for subagent work, it costs $0.10/$0.50 per million input/output tokens on prompts up to 100K (90% less than Haiku 4.5). Also, Claude's Python and TypeScript SDKs now handle computer and browser use in beta, saving you from writing a loop to pass Claude's clicks to drivers like Browserbase. Docs here.
Google turns text prompts into playable browser games: The search giant just unveiled Playground, a Google Labs experiment that lets you describe a 2D or 3D game in chat, then tweak its physics and rules as you go. Teams can quickly try out ideas before putting engineering time behind them. Google also announced a coming integration with Unity Spark, now in testing, to turn those prototypes into full Unity builds. Playground is live for U.S. users and you can try it here.

PRESENTED BY SOFTR
Your teams are vibe coding. How does IT stay in control?
Instead of blocking teams, give them a governed place to build with Softr.
IT defines guardrails and oversees Softr apps, connected data, users, and permissions.
Softr enforces access controls outside AI-generated code. Credentials stay server-side.
SSO, workspace roles, and audit logs are built in.
Softr handles authentication, permissions, hosting, and secure data connections.

INSIGHT
The eval playbook for anything you ship on an LLM

Source: The Code, Superhuman
Back it up. Your AI nailed the demo, but then you tweaked the prompt, and now you're not sure whether it got better or just sounds better. That's where evals come in, they’re basically repeatable tests that help you decide what's ready to ship. Senior engineer Sergii Makarevych's new guide walks through how to build them, catch convincing mistakes, and check whether your tests deserve your trust.
One good answer proves nothing. Makarevych breaks down how it works using an example of a runnable support bot for a fictional bike shop:
Build the loop. Collect real inputs and known-good answers, run your system through them, and grade every reply. Report a score with an error bar, then compare it with the previous version. Now you have something firmer than "this prompt feels better."
Read before you grade. Only 34 of the bot's 60 test replies passed. Cheap code checks caught missing facts but let an invented refund window slip through. A judge model can catch that, but it needs checking too. This one scored just 0.15 on kappa, a measure of agreement with human labels, well below the usual 0.6 release bar.
Compare on the same cases. Two prompts both scored 70%, but the new one did worse on 14% of tickets. With 60 cases, you can't reliably detect a drop smaller than about 12 percentage points. Check each pipeline stage separately. For agents, repeat tasks and track consistency, path, and cost per correct outcome.
Benchmarks pick the engine. A leaderboard can help you shortlist models, but scores shift with the setup and age quickly. Your own evals should decide what ships, with a clear threshold for blocking changes and monitoring once they're live. Makarevych's full essay digs into judge biases, agent reliability, and tooling. Pair it with former GitHub engineer Hamel Husain's Your AI Product Needs Evals for a real-product walkthrough.

PRESENTED BY WORKOS
AuthKit is the complete enterprise auth solution for your app, with user management at no cost up to 1M monthly active users.
Give enterprise customers the auth features they expect without building them from scratch. With AuthKit, you can:
Offer multiple auth methods including SSO, MFA, magic auth, & more
Add and remove users automatically with SCIM provisioning
Track who did what with audit logs
Ready to close your next enterprise customer?

IN THE KNOW
What’s trending on socials and headlines

Meme of the day.
Testing Tests: A senior engineer claims AI-written tests are empirically useless after his agent solved more tasks without them. Here's what he says to tell your agents instead (1.3K bookmarks).
Local Copilot: Microsoft CEO Satya Nadella says every Windows PC is now "an infinite software factory." Here's everything they launched to back it up (6.3K likes).
Drive Wipe: A founder says Claude Code wiped his C drive while clearing temp folders. A tiny mistake caused it, and Anthropic responded with how to prevent it (5.4K likes).
Smart Home: Take a look at the new smart home gear Apple and LG teamed up on (1.7K likes).
Fridge Phone: A dev turned Apple's foldable iPhone Duo into a working mini-fridge demo. It's giving 2008 joke-app energy, and people want more (151K views).
Clean Desktop: If your desktop is buried in screenshots, this open-source Mac app pins them to a clothesline on your screen, complete with drag-and-drop and one-click copying (12.5K likes).

TOP & TRENDING RESOURCES
Top Tutorial
How to build your own AI software factory with Claude Code: This tutorial walks through building an agent workflow that can take a feature from task to implementation and approval. You’ll also see how to run multiple features in parallel and extend the setup into a bigger software factory.
Top Repo
Agent Scripts (7.3k ⭐): From OpenClaw creator Peter Steinberger, this repo packages the shared rules, reusable skills, and lightweight helpers he uses across his agent workflows. It keeps Claude Code and Codex working from the same instructions.
Trending Cookbook
When to load MCP tools (by Google): This guide explains whether your agent should load every MCP tool upfront or only fetch tools when it needs them. It breaks down the trade-offs, plus when it makes sense to hand tool use off to a subagent instead.

AI CODING HACK
How to catch the risks buried in Claude Code's output
Claude Code ends every task with a wall of summary text, and the one line that matters gets skimmed past. The Claude Code team shared a built-in plugin called “You Should Know” that flags it for you.
Step 1: Enable it with one command inside Claude Code:
/plugin enable cc-plugin-you-should-know@builtinStep 2: That's the whole setup. The plugin scans Claude's output for information you might miss and surfaces it as a "Heads up" callout.
P.S. Get 50+ AI coding hacks for Claude Code, Cursor, and Codex here.

IN CASE YOU MISSED IT
Our most-clicked story from yesterday
Claude Code's creator shared how he actually prompts Claude. He says three things now matter more than the prompt itself.
Grow customers & revenue: Join companies like Google, IBM, and Datadog. Showcase your product to our 350K+ engineers and 150K+ followers on socials. Get in touch.
Whenever you're ready to dive deeper
We put together a few guides on coding agents, agentic engineering, and leadership frameworks to help you level up in your career. Browse all our guides.
What did you think of today's newsletter?
You can also reply directly to this email if you have suggestions, feedback, or questions.
Until next time — The Code team





