Editorial window: August 8 through August 14, 2026, America/New_York Audio rendered: August 29, 2026
Series note: AI Tech Signals is Triangle Technology Signals’ methodology podcast for technology decision-makers. Its operating guidance can apply beyond the Triangle; an episode counts as people-centered local reporting only when it names and sources local actors.
Central argument
This week’s evidence points to a practical caution: a business builder should define outcomes and proof before copying a human engineering ritual wholesale into the agent loop. More reliable comparisons need repeated evidence and a frozen model configuration, especially when defaults and behavior are changing.
Chapters
- 00:00: Opening argument
- 00:26: Direct the outcome, not the ritual
- 03:19: One convincing run is weak evidence
- 05:21: Freeze the setup before comparing models
- 07:04: Give each agent a distinct responsibility
- 08:30: Closing thought
Sources
- Birgitta Böckeler on MartinFowler.com: TDD inside the agent loop
- Birgitta Böckeler experiment repository
- Prime Intellect: Measuring Autonomous AI Research
- OpenAI Agents SDK JavaScript 0.15.0
- OpenAI Agents SDK JavaScript 0.16.0
- Google Antigravity changelog
- Google: Gemini 3.7 Flash
- Google: Introducing Custom Agents
- GitHub: Agent Plugins 1.0
Evidence cautions
- The TDD comparison was a small exploratory study and used model judgment for some quality dimensions.
- Prime Intellect’s benchmark was specialized, expensive, and noisy.
- Google benchmark claims and customer quotations in the source announcement are vendor-reported evidence.
- Feature availability does not prove that adding agents or plugins improves a specific tool.
