The Pulse — March 5, 2026
The signals that entered our radar, organized with sources and context to understand what changed.
The audio script is ready; narration will appear after voice generation finishes.
Gemini 3.1 Flash‑Lite (preview) — “intelligence at scale” pricing + speed
WHY IT ENTERED THE RADARGoogle is explicitly optimizing for high-volume production workloads (translation, moderation, UI generation, simulations) with very aggressive pricing ($0.25/M input, $1.50/M output) and better latency than 2.5 Flash.
SUGGESTED EDITORIAL ANGLEOpen original source ↗“The real battle isn’t ‘best model’, it’s best model per dollar per second—and Google is pushing that hard.”
Gemini 3.1 Flash‑Lite model card: the numbers (speed, context, evals)
WHY IT ENTERED THE RADARIt’s one of the rare “official” sources that publishes output speed (tokens/s) and a broad eval matrix. Also confirms 1M context window and 64K output.
SUGGESTED EDITORIAL ANGLEOpen original source ↗“Stop arguing vibes—here’s what the model card actually says: speed, cost, and where it wins/loses.”
Poetiq’s ARC‑AGI‑2 solver verified: 54% on semi‑private test (and open‑source)
WHY IT ENTERED THE RADARThis is the clearest “reasoning harness beats raw model” story: Poetiq reports ARC Prize verification of 54% at ~$30.57/problem, beating a previous best (they cite 45% for Gemini 3 Deep Think).
SUGGESTED EDITORIAL ANGLEOpen original source ↗“The new fine-tuning is test-time systems engineering: refinement loops, verification, and meta-optimization.”
Full‑duplex speech‑to‑speech on Apple Silicon: PersonaPlex 7B + Swift/MLX
WHY IT ENTERED THE RADARA practical “voice agent” stack is emerging without the classic ASR→LLM→TTS pipeline. This is closer to the UX people expect (interruptible, streaming, low-latency), and it runs locally.
SUGGESTED EDITORIAL ANGLEOpen original source ↗“Voice agents are about to feel alive: one model listening + speaking in real time—on a laptop.”
StepFun Step‑3.5‑Flash (open MoE) + paper + “agentic” cookbooks
WHY IT ENTERED THE RADARAnother strong signal that open models are being packaged as agent platforms (cookbooks for OpenClaw / Claude Code / local agents), not just weights. Also highlights MoE efficiency claims (activate ~11B of 196B per token).
SUGGESTED EDITORIAL ANGLEOpen original source ↗“Open models are copying the closed playbook: not just a model—an agent ecosystem with recipes.”
Qwen turbulence: leadership departures right after a strong open-weight run
WHY IT ENTERED THE RADARQwen 3.5 has been a cornerstone for “local + capable” setups. If the team destabilizes, it can ripple through open tooling, fine-tunes, and on-device apps.
SUGGESTED EDITORIAL ANGLEOpen original source ↗“Open weight momentum is fragile: one team ships half the ecosystem, and a re-org can change everything.”
AI-assisted rewrite to relicense (chardet): a real legal test case for ‘clean-room via LLM’
WHY IT ENTERED THE RADARThis is upstream of a huge future fight: can you use AI to “rewrite” copyleft code and then claim it’s clean-room? The post lays out the derivative-work trap and the human-authorship trap.
SUGGESTED EDITORIAL ANGLEOpen original source ↗“AI codegen might accidentally nuke the concept of relicensing—and maybe even copyleft itself.”
Matt Wolfe
WHY IT ENTERED THE RADAROn-device AI is shifting from “dev hobby” to consumer UX (Siri integration, shortcuts, voice mode). This is where mass adoption can happen.
SUGGESTED EDITORIAL ANGLEOpen original source ↗“Local AI is no longer ‘Terminal stuff’: it’s shipping as an App Store product with real UX.”
Y Combinator
WHY IT ENTERED THE RADAR“Reasoning harnesses” and refinement loops are becoming a mainstream narrative—worth jumping on early.
SUGGESTED EDITORIAL ANGLEOpen original source ↗“Fine-tuning is ‘the old way’ for many tasks; verification + refinement loops are the new meta.”