THE AI PULSEEN

The Pulse — July 28, 2026

The signals that entered our radar, organized with sources and context to understand what changed.

ModelsAgentsAnthropic
LISTEN TO THIS EDITION
The Pulse — July 28, 2026
Now in the PulseThis edition’s editorial thesis
RegusciLabs Pulse
Get the next Pulse directly.
  1. 01Kimi technical blog (primary)

    Kimi K3 releases open 2.8T-class weights

    WHY IT ENTERED THE RADAR

    Kimi says K3 is a 2.8T-parameter, 16-of-896-expert MoE with native vision and a 1M-token context window; the weights were promised by July 27. Its claims—long-running coding, compiler/kernel work, research agents, and video work—make this a genuine “open frontier” moment, not merely another cheap model.

    SUGGESTED EDITORIAL ANGLE

    “Open source just crossed 3 trillion parameters—what can you actually run or use?” Separate the headline parameter count from active parameters, hardware reality, API availability, and independently verified results.

    Open original source ↗
  2. 02Google (primary, Jul 21)

    Gemini 3.6 Flash: the agent economics story

    WHY IT ENTERED THE RADAR

    Google positions 3.6 Flash as a cheaper agent workhorse: $1.50/M input and $7.50/M output tokens, 17% fewer output tokens than 3.5 Flash, and reported gains on coding, computer use, and research benchmarks. Flash-Lite targets high-throughput subagents at 350 output tokens/sec.

    SUGGESTED EDITORIAL ANGLE

    “The model that makes agents cheaper by writing less.” Show why reducing tool calls and output tokens can beat marginal benchmark gains for a production workflow.

    Open original source ↗
  3. 03FermiSense case study (primary, Jul 27)

    A $500 RL fine-tune of a 9B model reportedly beats frontier models on one real task

    WHY IT ENTERED THE RADAR

    On its catalog-review workflow, FermiSense reports a GRPO-tuned 9B open model scored 87.3% versus 76.9% for its best frontier configuration, at $0.50 per 1,000 listings. That is a powerful case for treating frontier models as a data-generation baseline—not the final production system.

    SUGGESTED EDITORIAL ANGLE

    “A tiny fine-tuned model beat the frontier—but don’t copy the headline.” Explain the conditions required: a narrow task, a reliable scorer/eval, proprietary examples, and enough volume to justify training.

    Open original source ↗
  4. 04Anthropic / Dario Amodei (primary, Jul 27)

    Anthropic’s open-weights position: no ban, but mandatory capability testing

    WHY IT ENTERED THE RADAR

    Anthropic explicitly rejects a blanket ban on open weights. It advocates chip controls, action on industrial-scale distillation, and mandatory safety testing for sufficiently capable models—open and closed. This is a nuanced policy story amid the Kimi release.

    SUGGESTED EDITORIAL ANGLE

    “Anthropic is not asking to ban open models. Here’s what it is asking for.” Use the three-policy framework and contrast it with the social-media framing.

    Open original source ↗
  5. 05Claude Code changelog (primary)

    Claude Code’s new default: Opus 5 with 1M context—and nested agent teams

    WHY IT ENTERED THE RADAR

    Version 2.1.220 makes Opus 5 the default Opus model with a 1M context window and adds strict network allowlists, a DirectoryAdded hook, and forwarding for nested subagents. The product story is not just model IQ: it is safer, more observable multi-agent execution.

    SUGGESTED EDITORIAL ANGLE

    “The real Claude Code update isn’t the model—it’s agent-team plumbing.” Demo the mental model: coordinator → subagents → nested specialists, then explain why permissions and network constraints matter.

    Open original source ↗
  6. 06HumanLayer’s SlopCodeBench experiment (primary analysis)

    Better coding agents still struggle to preserve a codebase over time

    WHY IT ENTERED THE RADAR

    In a small 17-checkpoint SlopCodeBench subset, HumanLayer reports Opus 5 at 4/17 strict passes (24%); no model reached the end of any challenge with everything passing. The benchmark reveals a gap ordinary “one-shot issue” benchmarks hide: handling evolving requirements without accumulating regressions.

    SUGGESTED EDITORIAL ANGLE

    “Why coding-agent demos lie (a little).” Contrast a spectacular first task with the less glamorous test: can the agent change the project ten times without breaking old behavior?

    Open original source ↗
  7. 07AI Builder Club’s Open Agent Teams skill (upstream repo)

    Open-agent orchestration is becoming a reusable workflow, not a bespoke build

    WHY IT ENTERED THE RADAR

    The project packages a pragmatic pattern for dispatching any CLI agent in tmux with file-based completion signals and multi-turn iteration. AI Jason’s July 20 upload referenced it while arguing that persistent “sidekicks” and orchestration reduce context/token waste.

    SUGGESTED EDITORIAL ANGLE

    “Your agent team needs a completion protocol, not more prompts.” Explain the unsexy reliability layer—observable workers, durable handoffs, and a coordinator that can retry.

    Open original source ↗
  8. 08OpenAI research (primary, Jul 27)

    OpenAI finds that AI use is already reshuffling tasks across job titles

    WHY IT ENTERED THE RADAR

    From 800,000+ U.S. ChatGPT messages, OpenAI reports that 43.5% of occupation-specific work messages concern tasks associated with another occupation. Customer-experience, design, HR, legal, and marketing workers show especially high “task crossover.”

    SUGGESTED EDITORIAL ANGLE

    “AI may not replace your job title—it may steal your next handoff.” Frame it around one person doing the first 80% of a task that used to require legal, data, design, or engineering help.

    Open original source ↗
TAKE THIS PULSE TO YOUR AI

Continue the analysis where you already work.

Copy this prompt into ChatGPT, Claude, Gemini, or whichever AI you use. It includes the signals, sources, and a guide for turning them into decisions.

No account is connected and no data is shared automatically.
PROMPT.md