In partnership with

Good morning. It’s Friday, July 17th.

The big story last night was Kimi K3, Moonshot AI’s massive 2.8-trillion-parameter open-weight model from China. It surged to the top of the Frontend Code Arena, surpassing Claude Fable 5, while posting a strong 88.3 on Terminal Bench 2.1, just behind GPT-5.6 Sol.

For heavy coding/agentic workloads (where output tokens dominate), Kimi K3 can be 2–3x cheaper than the top Western flagships while trading blows on performance, and once the open weights drop (expected July 27th), self-hosting or third-party inference could drive costs even lower.

Totally unexpected that an open weight model would score 1st on some of these coding leaderboards. Things are happening fast.

-Jeff
AI Breakfast

You read. We listen. Let us know what you think by replying to this email.

How can AI power your income?

Ready to transform artificial intelligence from a buzzword into your personal revenue generator

HubSpot’s groundbreaking guide "200+ AI-Powered Income Ideas" is your gateway to financial innovation in the digital age.

Inside you'll discover:

  • A curated collection of 200+ profitable opportunities spanning content creation, e-commerce, gaming, and emerging digital markets—each vetted for real-world potential

  • Step-by-step implementation guides designed for beginners, making AI accessible regardless of your technical background

  • Cutting-edge strategies aligned with current market trends, ensuring your ventures stay ahead of the curve

Download your guide today and unlock a future where artificial intelligence powers your success. Your next income stream is waiting.

Moonshot AI drops Kimi K3 as world's largest open weights model

Moonshot AI dropped Kimi K3, a 2.8-trillion parameter mixture-of-experts model. It lands July 27, 2026, as the largest open weights release in history, going toe-to-toe with proprietary giants like GPT-5.6 Sol and Claude Fable 5, completely shifting the narrative that open models always lag behind closed tech.

The architecture scales to 896 total experts but activates only 16 per request. To run a 1-million-token context window, it couples Kimi Delta Attention, yielding 6.3x faster decoding, with Attention Residuals, per-head Muon optimization, and quantile load balancing. Native visual feedback loops allow autonomous agents to write frontend code, audit screenshots, and self-correct instantly.

Data proves the scale. K3 jumped 17 places to dethrone Claude Fable 5 in the web development arena with a 76% pairwise win rate. It captured the number two spot globally on the Vals Index, outperforming Opus 4.8. Artificial Analysis benchmarks confirm these frontier reasoning capabilities, though K3 shows an elevated hallucination rate, and ProgramBench creators note its metrics give generous credit for partial completions.

The era of free compute is over. At $3 input and $15 output per million tokens, Moonshot establishes a premium pricing tier. Yet, it remains highly cost-effective compared to Western models, squeezing Silicon Valley margins and rewriting the geopolitical race for self-hosted infrastructure. A Hong Kong IPO is next.

OpenAI’s first physical hardware is a $230 macro pad for managing AI agents

OpenAI partnered with keyboard company Work Louder to drop the Codex Micro, a $230 physical macro pad built for managing AI agents. Instead of typing prompts, developers coordinate workflows via programmable keys, dual joysticks, a rotary dial, and six RGB status lights indicating agent state. Configured via Work Louder's Input software, the Bluetooth and USB-C controller operates across six layers to control debugging and code reviews.

While OpenAI offers tactile keys to oversee these agent threads, it has locked down the backend. Flagship models like GPT-5.6 Sol and Terra now encrypt agent-to-agent communication, making it impossible for developers to audit how subagents delegate tasks. This security posture matches a massive capacity expansion: OpenAI lifted rate limits on GPT-5.6 Sol for its eight million active users. The flagship model recently cleared a 30-year-old statistics conjecture regarding the Benjamini-Hochberg procedure in 90 minutes.

To secure these systems, OpenAI runs GPT-Red, an internal model trained through self-play reinforcement learning to attack other AIs. GPT-Red locates vulnerabilities in 84% of novel scenarios, vastly outperforming the 13% human rate, making Sol six times more resilient to prompt injections.

Yet local runtime risks remain high. If sandbox protections are disabled, GPT-5.6 Codex has accidentally deleted home directories on macOS and Linux by incorrectly managing the $HOME environment variable during temporary directory creation.

Anthropic's Boris Cherny urges engineers to code rules directly for AI agents

Anthropic’s Boris Cherny is telling engineering teams to stop treating AI agents like temporary contractors and start embedding institutional rules directly into their codebases. Instead of letting agents guess, teams are using lint rules, CI pipelines, and dedicated files like CLAUDE.md to set hard guardrails. This cuts down on wasted tokens, automates annoying code reviews, and lets non-developers ship code safely.

This focus on developer infrastructure is paying off.

Anthropic is reportedly working with Goldman Sachs, Morgan Stanley, and JPMorgan to host investor meetings for a potential IPO as early as October. Fresh off a 965 billion dollar valuation, the Claude creator could beat OpenAI to the public markets. For investors, this is the first real chance to buy directly into enterprise AI developer tools like Claude Code, proving that the real value in this boom is moving toward structured, automated workflows.

Frontier models and product moves

Agents and the agentic stack

Business, labor, and institutions

Security and surveillance

Hardware and infrastructure

Research and science

Robotics

Media and creative AI

Watch

NameThatUI maps UI descriptions to precise developer terms, helping you prompt AI coding agents with perfect accuracy.

In Parallel MCP server exposes live company context and plan state to any AI agent via permission-scoped workspace URLs.

Velo 3.0 uses MCP and company knowledge to auto-generate, voice-clone, and localize product videos from screen recordings or prompts.

V2Fun creates 3D characters with 8K textures and AI motion capture from prompts, images, or videos instantly.

Graft AI maps legacy apps and screen-trapped workflows into stable agent tools that automatically self-repair when user interfaces change.

Thank you for reading today’s edition.

Your feedback is valuable. Respond to this email and tell us how you think we could add more value to this newsletter.

Interested in reaching smart readers like you? To become an AI Breakfast sponsor, reply to this email or DM us on X!

Thinking of starting your own newsletter? AI Breakfast readers who sign up with Beehiiv receive a 14-day free trial and 20% off for 3 months.

Keep Reading