When Anthropic published Building Effective Agents in late 2024, it did something rare for the industry: it gave engineers a shared vocabulary. Instead of another manifesto about artificial general intelligence, the guide offered six clear patterns for structuring LLM systems. A year and a half later, in 2026, the landscape looks radically different. Model Context Protocol has become a universal standard. Claude has gained new capabilities. Most organizations now have at least one agent running in production. Against that backdrop, it is fair to ask whether those six patterns still matter, or if they belong in the archive next to last year’s model weights.

I tested all six patterns against a local model in a side repository to find out. The answer is yes. They still hold up. But not because they are immutable laws. They hold up because the last eighteen months of production experience have validated the framework’s core logic.

What the Framework Actually Gave Us

The six patterns are worth remembering precisely: Prompt Chaining, Routing, Parallelization, Evaluator-Optimizer, Orchestrator-Workers, and Autonomous Agents. That last one is essentially a loop in which the model plans, acts, observes, and repeats until some condition is met.

Plenty of engineers were already chaining prompts or delegating tasks to worker threads before the guide appeared. What Anthropic provided was taxonomy. One person’s “agent” was another person’s “workflow,” and a third person’s “multi-step tool call.” The guide sorted the mess into buckets with clear boundaries. That made it possible to argue about trade-offs without talking past each other. In a field drowning in hype, crisp language is a kind of infrastructure.

The Industry Built On Top, Not Around

By 2026, these categories are baked into how teams design systems. Anthropic still teaches them in their Academy courses. Research papers and engineering blogs still use the same six buckets to describe new architectures. That kind of longevity is unusual for a discipline that refreshes its stack every quarter.

The reason is straightforward. The industry did not replace the framework. It built on top of it. New tools like MCP and the newer Agent Skills standards function as plumbing. They make it easier to connect a model to a database, expose a tool, or manage state. But they do not change the logic of when to use a router instead of an orchestrator. A better pipe does not rewrite the floor plan.

Production data in 2026 confirms this. The most common deployment pattern is still a single tool-use call paired with human review. The second most common is a multi-step workflow with exactly one handoff to a person. Both are direct descendants of Prompt Chaining and Routing. Full autonomous loops remain the exception, not the rule, in live systems.

Restraint Won the Market

The original guide’s best advice was also the advice most often ignored in 2024: use the simplest pattern that works. Do not deploy a full autonomous agent if a hardcoded path will get the job done.

The market has finally internalized this. Most agent pilots still fail, and they fail for the same predictable reason. Teams stack abstraction on abstraction until no one can trace the decision boundary. When the system drifts, debugging becomes archaeology. The companies that have succeeded in production are the ones that showed restraint. They defaulted to single-turn tool use. They added a routing layer only after the single prompt proved inconsistent. They treated autonomy as a liability to be justified, not a feature to be celebrated.

This is not an argument against ambition. It is an argument for composition. The patterns work best when you combine them deliberately rather than reflexively reaching for the most complex option on the menu.

Where the Seams Start to Leak

The framework is not a cure-all. There are hard limits that show up the moment you leave the prototype stage.

Para tareas de alta frecuencia y bajo costo, el código determinista sigue ganando. Un LLM no debería estar normalizando una columna de un CSV cuando pandas puede hacerlo en milisegundos sin alucinar. Evite los bucles autónomos si no puede definir un objetivo de evaluación preciso. Sin una condición de parada clara, el modelo iterará hasta que invente una razón para detenerse. Para decisiones de alto riesgo que requieren fundamentación externa, no dependa únicamente del conocimiento interno del modelo. Y vigile los cuellos de botella en la recuperación de datos. Cualquier patrón que dependa de la búsqueda vectorial o de APIs externas puede colapsar si su base de datos es lenta o si su ventana de contexto está saturada con fragmentos irrelevantes.

Estos no son casos límite hipotéticos. Son las limitaciones que separan una demo funcional de un sistema que sobrevive al fin de semana.

Una comprobación rígida y un fallo erróneo

Aprendí el valor práctico de este marco de trabajo mientras construía mi repositorio de pruebas. Estaba implementando el patrón Evaluator-Optimizer. Mi evaluador comenzó como una regex codificada que escaneaba la salida del modelo en busca de palabras clave específicas. El modelo devolvió una respuesta correcta y bien razonada que, casualmente, utilizaba sinónimos en lugar de las palabras exactas que yo buscaba. El evaluador lo marcó como un fallo.

El modelo tenía razón. Mi comprobación era demasiado rígida.

Solucionarlo requirió más que simplemente ampliar una lista de palabras. Cambié el evaluador mismo por un juicio basado en un LLM. Eso costó tokens adicionales y unos cuantos milisegundos más, pero devolvió la evaluación al nivel de abstracción adecuado. El patrón en sí era sólido. Simplemente había elegido la implementación incorrecta para la tarea. Ese es exactamente el tipo de error que el marco de trabajo pretende prevenir. Algunas evaluaciones necesitan código. Otras necesitan un modelo. Saber cuál es cuál es el objetivo principal.

Cómo utilizarlos ahora

Trate estos seis patrones como un punto de partida, no como una ley absoluta. Comience con un único prompt. Si la calidad es inconsistente entre los tipos de entrada, añada una capa de enrutamiento para enviar diferentes solicitudes a prompts especializados. Si necesita múltiples perspectivas independientes antes de tomar una decisión, utilice Parallelization. Si la tarea es grande y divisible, pruebe con Orchestrator-Workers. Solo recurra al bucle autónomo completo cuando el espacio del problema sea demasiado amplio para mapearlo previamente y cuando tenga un confiable