Agent Engineering

The practical craft of building agents that survive contact with real work: tool design, context, evaluation, failure modes and the boundaries worth enforcing.

A compact clockwork greenhouse tends one thriving plant while an enormous unfinished autonomous factory stands idle in the distance
Build the Smallest Autonomous Loop That Can Work
A practical way to scope AI autonomy around one observable goal, one bounded action surface, and one trustworthy feedback loop before expanding further.
Published on
A bridge section built by one industrial rig is subjected to independent alignment and load tests by a separate inspection system
Verification Must Be Independent
Why agents should not be the sole judges of their own work, and how separate evidence, evaluators, and invariants create trustworthy completion.
Published on
A powerful experimental machine operates inside a transparent containment chamber with only a few permitted tool channels crossing concentric safety walls
Sandboxes Are Capability Containers
Why effective AI sandboxing is about constraining authority, data, tools, networks, and persistence—not merely isolating a process.
Published on
Turbulent blue material passes through a precision industrial mold and emerges as exact illuminated forms ready for a downstream machine
Structured Outputs Are Protocol Boundaries
Why schemas, validation, repair, and semantic checks are the boundary that turns probabilistic model output into dependable system behavior.
Published on
A vast mechanical observatory directs most work through efficient instruments while reserving one luminous resource-intensive engine for a difficult observation
AI Cost Is an Architecture Problem
Why token budgets, model choice, context assembly, retries, and verification should be designed as part of the system rather than optimized after the bill arrives.
Published on
Rugged autonomous machines relay the same illuminated baton and field ledger across a stormy mountain route toward dawn
Durable Execution for Long-Running Agents
How checkpoints, leases, event history, reconciliation, and resumable steps let agent workflows survive the failures that real work inevitably encounters.
Published on
A precision instrument maker fits one exact brass coupling between a luminous core and a consequential machine while unused tools hang in shadow
Tool Contracts Are More Important Than Tool Count
Why reliable agents need narrow semantics, explicit effects, typed failures, and verifiable results more than they need an enormous catalog of tools.
Published on
A moonlit railway switchyard routes work among a precision instrument, a fast engine, a heavy locomotive, and a compact workshop car
Model Routing Is a Product Policy
Why choosing a model per task should express product risk, latency, privacy, and quality policy instead of hiding behind a benchmark leaderboard.
Published on
An accident investigation workshop reconstructing the illuminated path of an autonomous machine from scattered physical evidence and recorded signals
AI Observability Is Decision Reconstruction
How to trace context, model decisions, tool effects, policy, and evidence so agent behavior can be understood and improved after the run.
Published on
Page 1 of 4