Kimi K3 Review: The 2.8T Open Model That Beats Claude on Paper
Kimi K3 review after a week on real code: Moonshot's 2.8T open-weight model tops benchmarks at a third of Claude's price, but it's slow and hallucinates.
Kimi K3 review after a week on real code: Moonshot's 2.8T open-weight model tops benchmarks at a third of Claude's price, but it's slow and hallucinates.
Django-Bolt puts Actix and Rust under Django to hit 188K RPS. I dug through the source and benchmarks — where it beats FastAPI, and where the speed evaporates.
Claude Sonnet 5 review after a week: it nearly matches Opus 4.8 on coding, beats it on Terminal-Bench, and runs at just $2/$10 — but the tokenizer bites.
Bumblebee scans npm, PyPI, Go, MCP configs, and editor extensions for compromised packages, all without running a single install script. Hands-on review.
GLM-5.2 scores 62.1 on SWE-bench Pro vs GPT-5.5's 58.6, ships under MIT, and costs $1.40/M input tokens. Benchmarks, pricing, and the China data question.
GPT-5.5 hits 82.7% on Terminal-Bench and uses 72% fewer tokens than Claude — but loses SWE-Bench Pro to Opus 4.7. Seven weeks of real agentic use, reviewed.
Claude Fable 5 hits 80.3% SWE-bench Pro and 29.3% FrontierCode Diamond. It also costs 2x Opus 4.8, retains your data 30 days, and silently falls back.
Windsurf became Devin Desktop on June 2. Cascade dies July 1. Here's what the rebrand, Devin Local, and ACP support mean after a week with the new IDE.
GitHub Copilot switched to AI credits on June 1. Token math per model, real session costs, and whether your $10/month Pro plan still makes sense.
Google Jules queues coding tasks, runs them in a cloud VM, and opens PRs while you sleep. Free tier gives 15 tasks/day. Here's what worked and what didn't.