Stuff about Software Engineering

Tag: AI (Page 1 of 4)

Agent Orchestration Is a Capability, Not a Platform

We should not confuse agent orchestration capabilities with an Agent Orchestrator platform. The enduring architecture is simpler: expose enterprise capabilities through thin, governed APIs; let authorized AI harnesses discover and compose them; and treat the harness as the Agent Runtime rather than as a mandatory central platform.

That means:

  • REST/OpenAPI for service contracts

  • OAuth/OIDC for identity and authorization

  • MCP or equivalent open mechanisms for machine-readable discovery where useful

  • the harness for execution, skills, scheduling, state, policy and orchestration

The harness might be a coding assistant, a local application, a cloud service, a domain-specific runtime or a custom application. The capability should survive replacement of any one of them.

I’m not disputing the need for orchestration capabilities; I’m disputing the assumption that those capabilities imply an enterprise Agent Orchestrator platform.

The market framing

The market is converging around the idea that enterprises need an orchestration or control layer for AI agents. Gartner describes a fragmented market in which BOAT, BPA, AI agent platforms, iPaaS and open-source frameworks can all provide agent orchestration, and frames the emerging competition as one over control of enterprise execution:

McKinsey uses the term agentic mesh for a composable, distributed and vendor-agnostic orchestration layer connecting agents and traditional systems. Its proposed capabilities include discovery, registries, authentication and authorization, observability, evaluations and governance. It also identifies repeatable and executable actions through secure APIs as foundational infrastructure for agentic systems:

There is clearly a real set of capabilities here. The questionable step is the move from capability to mandatory central platform.

Old wine in new bottles

This closely resembles the integration-platform discussion. Integration platforms and iPaaS products solved real problems, but the architectural mistake was allowing the product to become the architecture, effectively making all integration depend on one central platform.

The same pattern is emerging again around agent orchestration. Workflow execution, tool invocation, authentication, state, scheduling, retries, audit, observability, policy and human approval are all legitimate needs, but most are familiar software and integration capabilities. The genuinely new part is that an LLM can decide dynamically what to do next, and multiple autonomous runtimes can collaborate. That creates a need for orchestration capabilities; it does not automatically create a need for a new enterprise platform category.

The harness is the Agent Runtime

A harness gives an AI model the ability to do useful work by combining instructions, skills, tools, context, permissions, execution semantics and guardrails. Increasingly, harnesses also provide scheduling, state, evaluation and observability. In practical terms, that is already an Agent Runtime.

It might be Codex, another local runtime, a cloud-hosted service, a scientific runtime or a custom application. There is no architectural reason why all of these must sit beneath one central Agent Orchestrator. The runtime is replaceable; the enterprise capabilities it consumes should be enduring.

Great APIs are the architecture

The durable enterprise asset is the set of business and technical capabilities exposed through stable service boundaries. Applied to AI, the implication is simple: if a capability can be securely invoked through a governed API, any suitable harness can orchestrate it.

The enterprise standard can therefore remain thin:

  • standard web APIs, predominantly REST where appropriate

  • OpenAPI or equivalent machine-readable contracts

  • OAuth/OIDC authentication and authorization

  • clear scopes and least privilege

  • stable versioned contracts

  • consistent auditability and policy enforcement

If business systems, data services and document workflows expose capabilities this way, a suitably capable and authorized harness can use them. A local Codex runtime executing skills on a schedule can orchestrate work. So can a cloud service, a domain-specific runtime or a traditional application.

Discovery does not require a central orchestrator

A predictable objection is that, without a central Agent Orchestration Platform, a runtime needs another way to know which capabilities exist. That is a discovery problem, not an execution-centralization problem.

UDDI, WSDL, OData and OpenAPI all attempted to make services describable and discoverable. Their weakness was that a developer still had to understand the interface and wire it into an application. AI changes that equation because a capable runtime can interpret machine-readable descriptions and decide dynamically how to use the advertised capability.

As described in Old Wine, New Bottles: Why MCP Might Succeed Where UDDI and OData Failed, MCP can connect services to AI runtimes capable of interpreting discovered capabilities rather than merely cataloguing them.

The resulting pattern is straightforward: REST/OpenAPI describes what can be called, OAuth/OIDC governs who may call it, MCP or equivalent mechanisms can help machines discover it, and the harness decides how to compose and execute it. Registries and catalogs may still be useful, but they should describe capabilities rather than become mandatory execution paths.

Governance belongs at durable boundaries

This approach does not weaken governance; it moves governance to places that survive technology replacement. At the capability boundary, authentication, authorization, policy, auditing and data contracts should be enforced. At the runtime boundary, the harness should govern models, skills, tools, execution, observability and human approval.

That lets us change the harness, model provider or execution environment without redesigning enterprise capabilities.

The test

If we replace the agent orchestrator in five years, do we still possess the orchestration capability? If not, we have created a platform dependency rather than an enterprise capability.

Keep the capability when the platform changes

There is a legitimate need for orchestration: scheduling, tool invocation, state, security, policy, observability, evaluation, human approval and coordination. We should resist turning those needs into a mandatory enterprise Agent Orchestrator.

Enterprise capabilities should be exposed through thin, governed standards, while autonomous runtimes and harnesses remain free to discover and compose them wherever they run.

I’m not disputing the need for orchestration capabilities; I’m disputing the assumption that those capabilities imply an enterprise Agent Orchestrator platform.

Access to a Model Is Not a Working Capability

“We don’t need ChatGPT Enterprise. We already have Azure Foundry.”

“I can run a model locally, so my tokens are free.”

Both statements mistake access to one component for possession of the complete capability. When I evaluate an AI alternative, I want to know what work it enables us to complete, at what quality, with how much intervention, and at what total cost.

Getting a model to run is an engineering achievement. Making it a dependable working environment is a much larger undertaking.

Microsoft describes Foundry as a platform for building AI applications and agents. That makes it a legitimate option for constructing capabilities. It does not, by itself, establish that those capabilities already exist for the people who need to use them. Someone still has to assemble, integrate, evaluate, operate and maintain the resulting product.

That distinction is central to the case for complete AI products. The value includes the work users can perform immediately and the engineering they do not have to commission first.

The model and its environment work together

I have argued that the harness matters increasingly: the tools, context management, execution environment, skills and controls surrounding the model. But this should not be interpreted as “the model no longer matters.”

Models differ in their ability to plan, handle ambiguity, use tools and recover from mistakes. Training can explicitly develop these abilities. OpenAI’s Codex launch report describes reinforcement learning on real coding tasks across environments, including learning to run tests iteratively.

The model and its working environment therefore cannot be treated as entirely independent components. Reproducing a tool interface does not automatically reproduce the learned behaviour needed to use it effectively.

Research supports the importance of that interaction. SWE-agent showed that purpose-built interfaces improved agents’ ability to navigate repositories, edit code and execute tests.

In my own work, Codex and Claude Code have been substantially more effective than Copilot App with an arbitrarily selected model. That is my experience, not a universal benchmark. It is nevertheless why I assess the complete working system rather than assume that a model catalogue establishes equivalence.

Smaller models belong in the architecture

I already advocate using powerful models to define a solution and smaller models to implement clearly specified work.

In one of my own implementation experiments, a local model spent approximately an hour without producing a usable implementation. After a frontier model clarified the representation, algorithm and initial test sequence, the local model produced a working implementation. Tests and an independent reference comparison passed, although review still identified two small contract deviations.

The useful result was a division of labour. Frontier reasoning established direction; local execution performed much of the implementation.

Recent research describes a similar mechanism. The preprint Better Harnesses, Smaller Models found that adapting harnesses improved 16 of 21 task–model combinations, with seven closing the performance gap. Adaptation worked best for repetitive workflows and models with sufficient underlying capability.

Even Small Language Models are the Future of Agentic AI, a position paper strongly advocating smaller models, centres its argument on specialised, repetitive tasks and proposes heterogeneous systems where broader capabilities are required.

This is consistent with my approach: resolve ambiguity, capture decisions in durable artifacts, and delegate well-defined execution. It does not establish that every smaller model can independently perform the reasoning that made the delegation possible.

Local tokens still have a cost

Local inference can be economically attractive. Removing the provider’s token bill, however, leaves hardware, electricity, operation, integration and maintenance. It also leaves the time spent waiting, supervising, correcting and retrying.

Those costs should be measured on both sides. Hosted products also have limits, require review and can fail.

The relevant comparison is total cost per successfully completed and accepted task. AI Agents That Matter makes the underlying evaluation argument: accuracy and cost must be considered together, and model benchmarks should not be confused with evidence that an agent suits a particular application.

A cheap unsuccessful attempt is still an unsuccessful attempt. A more expensive system can be better value if it reliably returns useful capacity to its users.

Replaceability does not establish equivalence

This is a counterpoint to the previous pattern. Designing a system so that its model can be replaced does not establish that another model will perform the same work successfully.

My model-replaceability pattern argues that business rules, trusted data, workflows and evaluation criteria should remain under our control. We should be able to change the reasoning engine without surrendering the application’s intended behaviour.

Replaceability still requires a sufficiently capable replacement. It does not imply that every model, harness or product is equivalent. The mistake is treating the ability to swap a component as evidence that we have replaced the capability of the complete system.

I am open to local models, open weights, alternative vendors and systems built on Foundry. The decision should follow representative work: compare accepted results, elapsed time, human intervention and the full cost of providing the capability.

I’m open to replacing the vendor. Show me that we can replace the capability—and what it costs to do so.

Pattern: Make Frontier Models Replaceable

Making language models pluggable is useful, but it does not make an AI application independent of the model behind it.

A system may support several providers while still relying heavily on the behavior of one particular frontier model. When prompts carry implicit business rules, priorities and workflow decisions, the API can be replaceable while the business solution remains tightly coupled to a model.

This pattern applies primarily to business applications built on general-purpose frontier models. Specialized scientific models can be different: the model itself may embody a unique domain capability.

Treat the model as a replaceable reasoning engine

The durable business capability should live in the surrounding system: explicit business rules, authoritative data, tools, workflows, interface contracts, guardrails, evaluation criteria and deterministic processing.

The model reasons within those boundaries. It should not have to invent the boundaries each time it runs.

Consider an application that assesses requests for a refund. A model might interpret a customer’s explanation or identify missing information. Eligibility rules, approval limits and the steps required to issue a refund should be explicit. Replacing the model should not silently change the refund policy.

Use model upgrades as a diagnostic

Ask what happens if today’s model is replaced with a materially better one.

Better reasoning, instruction following, robustness and language quality are welcome. A better model may also handle cases that the previous model could not. Those improvements do not, by themselves, indicate an architectural problem.

The useful question is what changed. Does the system apply the same policy more reliably, or does it start making different policy choices? Does it understand difficult requests better, or does it change which requests deserve priority? Does it follow the workflow more accurately, or invent a different workflow?

Changes in business decisions, priorities or required steps are a reason to investigate hidden model dependence. An upgrade can expose how much application behavior was implicit in the prompt.

Recognize the prompt as hidden architecture

During prototyping, it is tempting to put everything into a prompt: how to interpret situations, what matters, which decisions to make, how to handle exceptions and what actions should follow.

This is fast. Over time, however, it can make business logic difficult to locate, test and govern. Model upgrades then become changes to application behavior, even when nobody intended to change the application.

The model has become both a reasoning engine and an implicit part of the application’s architecture.

Make the dependency explicit

  1. Specify business behavior. Define policies, priorities, decision rights and exception paths so that they can be reviewed independently of a model’s answers.
  2. Move stable logic outside inference. Use ordinary software for calculations, validation and rules that have a determinate answer.
  3. Ground reasoning in authoritative data. Give the model access to the relevant facts rather than relying on its general knowledge to supply business context.
  4. Put model calls behind clear contracts. Define inputs, expected outputs, validation and failure handling. Enforce permissions in the surrounding system.
  5. Keep prompts focused on the reasoning task. Make any policy instructions explicit and testable, and avoid relying on the model to fill gaps in the workflow.
  6. Evaluate across multiple capable models. Test important scenarios, including exceptions, and distinguish improvements in reasoning from changes in intended business behavior.

The aim is substantially stable business behavior when the reasoning engine changes. That still requires evaluation: a common interface cannot guarantee equivalent behavior across models.

Design for better models without depending on them

This pattern follows two broader engineering principles: use AI where ambiguity requires interpretation, and compose the application through explicit software contracts and workflows.

Those principles help determine where AI belongs. Replaceability asks how much the solution should depend on the particular model doing the reasoning.

Make the model easy to swap, but also make it possible to explain what must remain true after the swap. A better model should improve the system’s ability to perform its intended work. It should not be the first place where the intended work becomes defined.

Anti-Pattern: Using AI as a Runtime for Deterministic Workflows

AI can be remarkably effective at turning vague, manual work into software. But that does not mean an AI model should execute the resulting workflow forever.

A common anti-pattern is to use an AI agent as the runtime for stable business logic: reading the same kind of file, rediscovering its structure, applying known mappings, and generating the same kind of output on every run.

The first demo may look impressive. Repeated in production, it becomes expensive, fragile, difficult to test, and hard to govern.

Context

This pattern appears in repeatable work such as file transformation, CSV or Excel processing, data mapping, schema validation, reference-data lookup, and template configuration.

A typical workflow accepts a file and some known context—such as a site, market, customer, or product—then applies established rules, validates the result, and returns a configured output file.

These are primarily workflow and contract problems, not reasoning problems.

The Anti-Pattern

The anti-pattern is using an AI model as the execution engine for deterministic business logic.

Each run asks the model to rediscover the same structure, reinterpret the same rules, inspect the same reference data, and reconstruct the same transformation. The system consumes tokens as if the problem were new, even though the workflow is stable.

Stable, repeatable transformations should be implemented as deterministic workflows, not repeated AI reasoning sessions.

Why It Fails

  • High marginal cost: every run spends tokens on parsing, reasoning, and validation.
  • Poor determinism: the same input may not reliably produce the same output.
  • Hidden business logic: rules live in prompts and examples instead of reviewable code and configuration.
  • Weak validation: correctness depends on model behavior rather than explicit, testable contracts.
  • Limited control: failures are harder to reproduce, debug, audit, and explain.

The deeper problem is that the design confuses discovery with execution. A successful proof of concept is mistaken for a sustainable architecture.

The Better Pattern: Compile the Workflow

Separate design-time AI assistance from runtime execution.

At design time, use AI to:

  • understand the existing manual or spreadsheet-based workflow;
  • identify schemas, mappings, rules, and edge cases;
  • generate parser and transformation code;
  • create validation tests and documentation.

At runtime, use deterministic software to:

  • parse the input;
  • validate required fields and schemas;
  • look up governed reference data;
  • apply known mappings;
  • produce the output and a validation report.

The implementation might be a script, web application, serverless function, workflow step, or agent action. The form is secondary. The important part is deterministic execution behind a stable contract.

Example

Suppose a spreadsheet performs site-specific file configuration.

The wrong solution is an agent that repeatedly reads the spreadsheet, interprets reference files, reasons about the necessary changes, and generates a CSV.

The better solution is to reverse-engineer the logic once, define the input and output schemas, move reference data into governed configuration, implement the transformation in code, and add tests. An agent may still provide the conversational interface, but it should call the deterministic workflow rather than embody its logic.

When Runtime AI Is Appropriate

AI still belongs at runtime when the task requires genuine interpretation:

  • the input is novel or its structure changes constantly;
  • the output requires judgment, summarization, classification, or explanation;
  • the workflow is exploratory and has not stabilized;
  • codifying deterministic rules would cost more than repeatability is worth.

Once a workflow stabilizes, move it from repeated reasoning into deterministic execution.

Detection Signals

You may be using this anti-pattern if people say:

  • “The agent needs to understand the spreadsheet every time.”
  • “It works when I explain it carefully.”
  • “Sometimes it creates the file correctly.”
  • “We need more prompt guardrails.”
  • “It consumes many tokens, but the task is always the same.”

Rule of Thumb

If the same input, reference data, and context should always produce the same output, the runtime should be deterministic.

Use AI to build the machine, not to act as the machine.

The goal is not to spend tokens forever. It is to convert repeated reasoning into reusable capability.

The Bottleneck Moved

For years, the dominant constraint in software engineering was implementation cost. Writing software was expensive. Changing it was expensive. The whole discipline organized itself around that fact — estimation, sprints, backlogs, delivery teams — because throughput was the bottleneck.

AI is dismantling that bottleneck faster than most organizations have been able to think about what replaces it.

But there’s a less-noticed shift happening alongside the capability story. Quietly, and then quite visibly, the major AI vendors are moving coding workflows toward consumption-based pricing. GitHub Copilot has introduced premium requests. Anthropic has repeatedly adjusted Claude Code access as demand exploded. OpenAI has separated agent-style workflows into explicit credit models for enterprise use. The details differ — different token pools, different rate limits, different structures — but the direction is consistent: frontier inference is not economically unlimited, and the vendors are starting to price it that way.

This matters more than it might appear, for one reason: AI-assisted development is not reducing demand for software. It’s accelerating it. Developers produce more. Non-developers can now produce some. Agents execute implementation work in parallel. If software was already eating the world, AI is increasing the bite rate.

Which means total inference demand is likely to rise faster than efficiency gains bring costs down. Even as models get cheaper per token, the sheer volume of generated code, tests, pipelines, reports, and scaffolding will grow. Organizations routing large amounts of that work through frontier reasoning models are building a cost structure that may become uncomfortable quickly.


The current pattern of AI-assisted development tends to look like this: open a long-running session with a flagship model, iterate continuously, keep large context windows alive, regenerate as needed. It works. The problem is what it’s spending expensive reasoning capacity on.

A significant portion of that work is not actually ambiguous. It’s repetitive. Deterministic. Structurally predictable. Generating a stable HTML report structure for the fifteenth time doesn’t require the same model that helped you design the architecture. Implementing a transformation you’ve already fully specified doesn’t require frontier-level reasoning. But because the workflow isn’t designed to distinguish between the two, it routes everything through the same expensive path.

That’s partly an economic inefficiency. But it’s also a signal. If implementation continuously requires frontier-level reasoning, that’s usually not a model problem — it’s a decomposition problem. It means ambiguity that should have been resolved upstream is still alive in the work.

Historically, you could absorb that ambiguity because implementation throughput was the real constraint anyway. When implementation gets cheap, the constraint moves. What’s left is ambiguity, architectural clarity, decomposition quality, validation, and organizational alignment. That’s where the work now actually is.


The architectural implication is straightforward, even if acting on it isn’t. Use stronger models where ambiguity and interpretation are genuinely high — architecture, requirements, tradeoff analysis, decomposition, synthesis. Once that work is done and structure has been established, the remaining implementation should be deterministic enough that it doesn’t need the same reasoning capacity.

I’ve been exploring what this looks like in practice through two experimental repositories, Rupify and Speckify. Rupify concentrates expensive reasoning at the front of a workflow: AI-assisted stakeholder interviews, requirements normalization, formal specification generation. The goal is to resolve ambiguity deliberately and early, producing canonical artifacts that downstream work can execute against. Speckify then takes those specifications and decomposes them into atomic, traceable implementation units — work that’s structured enough that smaller, cheaper models can handle it reliably.

The point isn’t “better AI coding.” It’s redesigning the workflow so that frontier reasoning gets used once, where it’s genuinely needed, rather than continuously throughout the entire delivery cycle. The economics look very different when you do that.


For a while, the dominant question in AI-assisted engineering was capability: which model, which agent, which coding assistant. That question isn’t going away. But another constraint is now becoming visible alongside it.

The organizations that scale this well probably won’t be the ones that put AI everywhere. They’ll be the ones that figure out where ambiguity actually lives in their workflows, structure everything else aggressively, and stop paying frontier-model prices for work that stopped being uncertain three steps ago.

The competitive advantage isn’t more AI. It’s better placement.

Pattern: Efficient Use of AI for Software Engineering

Why this matters

AI-assisted development is moving to usage-based cost models where:

  • Cost scales with model choice and interaction patterns

  • Long-running sessions and iterative loops increase spend

  • Higher-capability models are significantly more expensive

At the same time, AI enables:

  • Rapid exploration of designs

  • Large-scale code generation

  • Parallel execution of implementation work

Without structure, this leads to:

  • Unpredictable cost

  • Inconsistent quality

  • Unnecessary rework

This pattern defines a structured way of working that maximizes value while controlling cost and complexity.


Core pattern

Use powerful models to define the solution. Use cheaper models to implement the solution.

The goal is to:

  • Concentrate reasoning effort once

  • Avoid repeated re-evaluation

  • Execute implementation in small, independent units


Step 1 — Use AI for solution design

Use the most capable model available to:

  • Explore solution options

  • Evaluate trade-offs

  • Validate architecture

  • Identify risks and edge cases

  • Define architecturally significant requirements (ASRs)

Output:

  • Solution design document

  • Clear constraints and assumptions

  • Agreed direction

This step should remove as much ambiguity as possible.


Step 2 — Make the solution executable

Translate the design into structured work:

  • Break into epics and issues

  • Define scope and expected outcome for each

  • Ensure each unit is:

    • Scoped

    • Unambiguous

    • Testable

Good decomposition is the primary control mechanism.

Well-defined units enable:

  • Predictable AI execution

  • Minimal context per task

  • Independent implementation


Step 3 — Execute with clean context

For each issue:

  • Start with a fresh context

  • Provide only:

    • The issue description

    • Relevant constraints

    • Local code context

Avoid:

  • Long-running chat sessions

  • Accumulated conversation history

  • Repeated "compaction" of context

Treat every task as a clean execution.


Step 4 — Use cheaper models for implementation

Once tasks are well-defined:

  • Use faster, lower-cost models

  • Focus on:

    • Implementation

    • Test generation

    • Applying patterns

If a task requires a high-end model:

The issue is likely underspecified.


Step 5 — Execute in parallel where possible

When issues are independent:

  • Use subagents or worktrees

  • Implement in parallel

  • Rely on:

    • Clear boundaries

    • Well-defined contracts

This enables:

  • Faster delivery

  • Better utilization of AI

  • Consistent results


Step 6 — Avoid long-context degradation

AI performance degrades in extended sessions:

  • Context grows

  • Signal-to-noise decreases

  • Output quality drops

Common anti-pattern:

  • Iterate continuously in one session

  • Compact context

  • Continue

This accumulates errors over time.

Recommended approach:

  • Keep interactions short

  • Reset frequently

  • Reintroduce clean inputs


Step 7 — Store state in artifacts

Do not rely on the model to maintain system state.

State should be captured in:

  • Design documents

  • Specifications

  • Epics and issues

This ensures:

  • Reproducibility

  • Consistency

  • Independence between tasks

The model executes — artifacts define the system.


Step 8 — Keep feedback loops controlled

Even with strong design:

  • Issues will evolve

  • Edge cases will appear

Handle this by:

  • Updating artifacts (not conversations)

  • Refining issues

  • Re-running tasks with clean context


Summary

This pattern enables:

  • Predictable cost

  • Consistent quality

  • Scalable execution

By:

  • Separating reasoning from implementation

  • Minimizing context

  • Structuring work into independent units

  • Executing with clean, repeatable inputs

Solve once. Structure clearly. Execute repeatedly.

When Implementation Becomes Cheap: Rethinking Value in Software Consulting

Introduction

There was a time when building software was the work.

Methods like Use Case Points (UCP) gave us a structured way to estimate implementation effort—because implementation was the dominant cost.

That assumption is now broken. AI coding agents have collapsed implementation time by an order of magnitude. At the same time, approaches like Rupify turn specifications into executable, verifiable inputs that can steer those agents [1].

This sets up a new tension.

The Real Tension: Speed vs. Correctness

The Wiggum Loop shows that you can brute-force progress with AI through rapid iteration [2]. But it also shows where that breaks: when failures are silent, slow, or irreversible. That is exactly the regime most enterprise systems operate in. So this is the core tension:

The Wiggum Loop is powerful precisely where it is dangerous. Fast iteration in domains where incorrect systems are costly. This is where Rupify resolves this tension. It does not slow the loop down. It constrains it.

It makes fast iteration safe enough to use in high-stakes environments by making intent explicit and verifiable.

The Failure Mode: Faster Divergence

Without that constraint, AI does not give you better outcomes. It just gives you incorrect systems, delivered quickly.

Organizationally, this looks like:

  • Systems that appear complete but encode the wrong logic
  • Silent misalignment between business intent and system behavior
  • Accelerated rework cycles where errors propagate faster than they are detected

The result is not efficiency. It is amplified waste. This is why specification becomes the control point.

The Inversion

This creates a structural inversion:

  • UCP still estimates human implementation effort
  • AI reduces actual implementation time to a fraction
  • Rupify ensures the output remains aligned with intent

So the traditional model—where effort ≈ implementation ≈ value—no longer holds.

UCP Was Measuring the Wrong Thing

You can still calculate UCP. But it no longer answers the original question: How long will this take to build?

That question is now nearly irrelevant, what remains is something more fundamental: How complex is the problem space?

UCP was always approximating this. So AI did not make UCP obsolete, it revealed what UCP was actually measuring all along.

Knowing complexity is crucial for coding with AI Agents [5].

A New Model

We end up with a new structure:

  • UCP measures problem complexity
  • Rupify translates complexity into executable intent [1]
  • AI agents handle implementation at near-zero marginal cost

What Consulting Becomes

This connects directly to outcome-based value models [4].

In this new model consulting is the discipline of reducing ambiguity. That is not a slogan. It is a structural shift.

It changes:

  • Staffing → fewer implementers, more domain modelers and specification engineers
  • Pricing → from time-based delivery to value of clarified and executable intent
  • Differentiation → ability to make complex systems unambiguous, not ability to build them

The scarce role is the person who can:

  • Extract intent from messy organizational reality
  • Structure it into a precise model
  • Express it in a form that machines can execute correctly

That capability becomes the bottleneck.

The Final Constraint: Can It Be Safely Realized?

Even perfect specifications are not sufficient.

They must be realized through a trustworthy system [3].

A correctly specified system built through a compromised pipeline is still a compromised system.

So the full model becomes:

  • Clarity of intent (Rupify)
  • Controllability of generation (AI agents)
  • Trustworthiness of realization (supply chain)

Remove any one of these, and the system fails.

The Shift

Software engineering is no longer about building systems.

It is about:

  • Describing them correctly
  • Constraining how they are generated
  • Ensuring they can be safely realized

Implementation has not disappeared, but it has lost its position as the center of value – and when that happens, everything around it has to be rethought.

Conclusion

Implementation is no longer the primary driver of cost, time, or value.

Value is created by reducing ambiguity, expressing intent precisely, and ensuring that intent can be safely realized.

  • AI accelerates execution
  • Rupify constrains it
  • UCP reveals the true complexity underneath

Consulting shifts from delivering software to making systems unambiguous and executable.

References

[1] https://birkholm-buch.dk/2026/04/09/rupify-executable-specifications-for-ai-assisted-software-engineering

[2] https://birkholm-buch.dk/2026/04/05/the-wiggum-loop-brute-forcing-business-with-ai/

[3] https://birkholm-buch.dk/2026/03/13/move-the-security-boundary-to-the-software-supply-chain/

[4] https://birkholm-buch.dk/2025/05/05/the-future-of-consulting-how-value-delivery-models-drive-better-client-outcomes/

[5] https://birkholm-buch.dk/2024/12/12/speed-vs-precision-in-ai-development/

AI Is Everywhere. Value Is Not (And It’s Not a Data Problem Either)

Introduction

Over the past year, AI adoption has exploded. In the Nordics, nearly every company now reports that it has implemented AI in some form. On paper, that should translate into a wave of productivity, growth, and competitive advantage. Only it doesn’t.

A recent BCG study (The Nordic AI Inflection Point: Value Creation or Value Bubble?) shows that while 99% of Nordic companies have adopted AI, only around 4% report significant returns on their investments. At the same time, executives expect AI to deliver 25–30% improvements in both revenue and cost.

This gap between adoption and value is not subtle and it’s not limited to the Nordics. A global enterprise study shows the same pattern (Enterprise AI adoption in 2026: Why 79% face challenges despite high investment):

  • Near-universal AI adoption
  • Heavy usage across employees and executives
  • Only a minority seeing real business impact

More strikingly, over half of executives report that AI adoption is creating internal tension rather than clarity — exposing gaps in strategy, ownership, and execution.

AI is not just failing quietly. It is actively stressing organizations that are not designed to absorb it. Which raises an uncomfortable question: Are we creating value — or a value bubble?

This Is Not a New Problem

In a previous post, I argued that AI doesn’t create advantage but distribution does based on facts that:

  • AI is becoming commoditized
  • Models are widely accessible
  • Tools are rapidly diffusing

So advantage cannot come from AI itself. It must come from how AI is embedded, scaled, and operationalized. The BCG findings are a direct confirmation of this. So AI is everywhere, but execution is not.

The Wrong Debate: Data Before AI

At the same time, many organizations seems to be stuck in a different discussion: “We need better data before we can scale AI.”

I’ve argued the opposite in “AI for data — not data before AI” and that waiting for perfect data is one of the most reliable ways to delay value indefinitely.

Data improves when it is used:

  • In real workflows
  • Under real decisions
  • With real feedback loops

So we end up with two truths:

  • AI alone does not create advantage
  • Data alone does not unlock AI

And yet, most organizations behave as if one of them will.

The Two Traps Killing AI Value

What we see in practice is a predictable pattern.

1. The Tool Trap

Companies deploy AI as tools:

  • Copilots
  • Assistants
  • Automation add-ons

These deliver local gains but they don’t change outcomes, they don’t scale and they don’t compound.

2. The Foundation Trap

Others go the opposite direction:

  • Multi-year data programs
  • Master data management initiatives
  • Platform modernization

AI becomes a future promise and not a present capability.

The False Choice

This leads to a false dichotomy:

  • AI first
  • Or data first

The reality is neither.

  • You don’t get better data before AI
  • You don’t get value from AI without execution

Both positions assume a linear path and AI value is not linear.

What Actually Works: AI in the Loop

The companies that are capturing real value are doing something different.

They are not thinking in steps like: Data → Platform → AI → Value

They are building feedback systems: AI → Usage → Better Data → Better Workflows → Scale → Value

But this only becomes real when you look at how it is designed.

A repeatable pattern looks like this:

  • Start with a concrete workflow (e.g. demand planning, pricing, campaign execution)
  • Apply AI to improve one critical decision point
  • Use the output to expose data gaps and inconsistencies
  • Fix only the data that matters for that workflow
  • Expand AI across adjacent steps
  • Gradually connect the process end-to-end

For example:

  • Deploy AI in demand forecasting
  • Uncover inconsistencies in product hierarchies and sales signals
  • Fix those selectively
  • Extend into inventory and replenishment

Over time, the workflow becomes:

  • More accurate
  • More automated
  • More integrated

This is not just iteration.

It is system design.

Good AI systems are not built top-down. They are grown through use — and then engineered for scale.

From Tools to Workflows

The BCG report highlights a critical distinction:

  • Most companies invest in tools
  • Leaders invest in workflows

That difference matters.

Because:

  • AI applied to tasks creates efficiency
  • AI embedded in workflows creates advantage

Why Most Companies Stall

When AI fails to scale, it’s rarely about the models.

It’s about the system.

  • Tool Trap → Fragmentation
  • Foundation Trap → Delay

Both lead to the same result:

  • Pilots everywhere
  • Duplication of effort
  • No compounding value

The deeper causes are structural:

  • Fragmented data
  • Decentralized ownership
  • Unclear decision rights
  • Limited execution capacity
  • AI treated as IT

The system is not designed to absorb and scale AI.

So AI remains additive and not transformative.

AI Doesn’t Fail. Systems Do.

AI is not underdelivering, but organizations are.

Or more precisely: AI doesn’t fail, it exposes systems that were already failing.

What we are seeing is not an AI gap, it’s a system gap:

  • Between ambition and execution
  • Between tools and transformation
  • Between experiments and scale

The Trilogy

Across three posts, the pattern becomes clear:

  1. AI doesn’t create advantage — distribution does
  2. AI for data — not data before AI
  3. AI is everywhere. Value is not

Together: AI value is not created by technology or data alone, it’s created by systems that connect them.

The Next Phase of AI

The next phase will not be defined by better models. It will be defined by better systems.

Today’s tools are built as standalone assistants:

  • Copilots
  • Chat interfaces
  • Isolated automation

They optimize individuals and not systems.

The tools themselves reinforce the Tool Trap. Which means: Organizations are not just using AI incorrectly and they are buying products that make correct usage harder.

What This Means in Practice

If you want to capture AI value:

  • Stop measuring progress by tools
  • Stop waiting for perfect data
  • Stop layering AI on top

Instead:

  • Start with workflows
  • Build feedback loops
  • Design for reuse
  • Treat AI as part of the operating model

This is not a maturity curve, it’s a design choice.

Conclusion

You don’t win with AI because you have access to it. You don’t win because your data is perfect. You win when your organization can turn AI into systems that scale.

Advantage comes from system design:

  • Not tools
  • Not data in isolation
  • Not default ways of working

Because in the end: AI doesn’t create advantage, distribution does – and distribution is built through systems — whether you design them intentionally or not.

The difference is simple: Some companies design them, but most don’t.

The Evolution of AI: From Frontier Models to Specialized Small Language Models

Where We Came From: The Frontier Model Plateau

Over the past 12–18 months, the large language model (LLM) ecosystem has continued to advance—but largely in an incremental, not disruptive, fashion. Models from OpenAI, Anthropic, and Google have steadily improved across reasoning, multimodality, and scientific benchmarks, yet the relative ordering and qualitative capabilities have remained broadly stable.

Public benchmark suites such as MMLU (Massive Multitask Language Understanding), GPQA (Graduate‑Level Google‑Proof Q&A), and HELM (Stanford Holistic Evaluation of Language Models) show year‑over‑year gains measured in percentage points rather than step‑function breakthroughs. This is not a criticism—these are remarkable systems—but it does indicate a phase of maturation rather than rupture. Frontier models are converging: better, more reliable, more general—but not fundamentally different.

For scientific research, this means frontier GenAI has become a dependable horizontal capability: excellent for literature synthesis, reasoning assistance, explanation, and orchestration—but no longer the sole locus of rapid innovation.

Where We Are Now: The Rise of Small and Specialized Models

In parallel, a very different dynamic is unfolding.

Small Language Models (SLMs) and domain‑specific foundation models are advancing rapidly, particularly in scientific domains such as genomics, protein science, chemistry, and materials research. These models fall broadly into two categories:

  1. Domain‑adapted language models – smaller LLMs fine‑tuned on specific scientific corpora (e.g. chemistry, biology, materials science).
  2. Non‑linguistic foundation models – transformer‑based models trained on alternative “languages” such as DNA, protein sequences, or molecular graphs (e.g. Evo2, ESM, AlphaFold‑class models).

These models are not generalists—and that is precisely their strength. They encode deep inductive bias for their domain, deliver strong signal from sparse data, and increasingly outperform general LLMs on narrowly scoped scientific tasks.

Critically, most of these models do not fit the SaaS GenAI paradigm. They are rarely available via Azure AI Foundry, AWS Bedrock, or similar managed services. Running them typically requires:

  • Dedicated GPU infrastructure (often NVIDIA‑specific)
  • Local fine‑tuning or adaptation
  • Tight coupling to data and experimental context

This creates a structural mismatch between where scientific model innovation is happening and where traditional enterprise AI platforms operate.

External Validation: SLMs as First-Class Scientific Tools

Recent academic work explicitly supports this shift toward small, specialized models. A 2025 paper, “SLMs as Scientific Tools” (arXiv:2512.15943), argues that capability in scientific AI is task-relative rather than size-relative. The authors show that domain-specialized SLMs can match or outperform frontier LLMs on constrained scientific tasks when correctness, structure, and tool integration matter more than linguistic breadth.

Several conclusions from the paper closely align with CRL’s direction:

  • Inference locality beats central intelligence: running models close to data improves latency, reproducibility, validation, and cost control—supporting local, HPC-adjacent, and desk-side deployment.
  • SLMs scale scientifically, not just economically: smaller models are easier to interpret, benchmark, and falsify—critical properties for hypothesis generation and experimental decision-making.
  • Tool integration matters more than prompt engineering: structured inputs and deterministic tool calls outperform free-form prompting in scientific workflows.

The paper ultimately reinforces a hybrid architectural stance: LLMs orchestrate; SLMs execute. This provides external, peer-reviewed validation that SLMs are not a compromise, but the correct abstraction for scientific computing.

A Practical Shift: From Cloud‑Only to Desk‑Side AI

This is where a meaningful, practical shift is occurring.

With the arrival of systems such as NVIDIA DGX Spark, small language models become physically accessible to individual researchers. Instead of renting over‑provisioned H100 or Grace‑Blackwell cloud instances, scientists can:

  • Run and fine‑tune SLMs locally
  • Experiment rapidly without cloud friction or cost surprises
  • Work directly with models that are otherwise unavailable as managed services

In effect, this enables a “small model on every scientist’s desk” paradigm. The value is not raw scale, but immediacy, ownership, and experimentation velocity.

At CRL, this aligns tightly with how scientific progress actually happens: iterative, exploratory, domain‑specific, and data‑proximate.

Looking Toward 2026: A Hybrid, Orchestrated Future

Looking ahead—without making speculative predictions—the most plausible trajectory is not LLMs versus SLMs, but LLMs plus SLMs.

A likely pattern is:

  • Frontier LLMs acting as generalist reasoning, planning, and orchestration layers
  • Specialized small models performing high‑fidelity domain work (genomics, proteins, chemistry, simulation)
  • Tool‑ and model‑calling as the primary integration mechanism

In this model, the LLM does not replace scientific models—it coordinates them. It becomes the interface and glue, while the real scientific signal is generated by specialized systems running locally or on targeted infrastructure.

This is not speculative technology. The building blocks already exist:

  • Tool‑calling and agent frameworks
  • Domain foundation models
  • Local GPU systems capable of running serious scientific workloads

What changes in 2026 is not the theory, but the accessibility.

Summary

  • Frontier LLMs are improving steadily, but incrementally
  • Scientific innovation is accelerating fastest in small, specialized models
  • These models do not fit cloud‑only GenAI platforms
  • Desk‑side systems like DGX Spark make SLMs practically accessible
  • The near‑term future is hybrid: generalist orchestration + specialist execution

Appendix: The Emerging Scientific SLM Ecosystem (snapshot as of 2026-01-21)

Vendor / OriginDomain FocusRepresentative ModelsTypical Scientific Use Cases
NVIDIABiology, Chemistry, ClimateBioNeMo, ChemGPT, MegaMolBART, FourCastNetMolecule generation, QSAR, virtual screening, protein design, weather & climate modeling
DeepMindHigh-impact scientific modelingAlphaFold 3, GraphCastProtein structure prediction, climate forecasting, large-scale simulation
MetaProteins, Scientific LiteratureESMFold, ProtBERT, SciBERTProtein folding, sequence modeling, scientific text analysis
Arc Institute / ProfluentDNA & Protein DesignEvo2, E1DNA sequence design, protein design, strain optimization
Academic & Research ConsortiaGenomics, Materials ScienceOpenFold, MaterialsBERT, MatSciBERTCrystal property prediction, materials discovery
Emerging VendorsSupply Chain & OptimizationSCGPT, Logistics-LLaMA, OR-LLMDemand forecasting, route optimization, constraint planning

Notes

  • Most models listed above are open, open‑weight, or research‑licensed, and evolve in close collaboration with the scientific community.
  • The ecosystem is interoperable and tool‑oriented, designed to be embedded into pipelines rather than accessed via chat interfaces.
  • In contrast, enterprise GenAI platforms primarily target closed, managed, productivity‑oriented workloads.
  • NVIDIA’s role is increasingly that of a horizontal scientific AI platform provider, spanning models, tooling, and local compute rather than acting as a single‑model vendor.
  • Unlike enterprise GenAI platforms, which are predominantly closed and productivity-oriented, the scientific SLM ecosystem is characterized by open models, research licensing, and composability— properties that align naturally with exploratory research environments such as CRL.

Rupify: Executable Specifications for AI-Assisted Software Engineering

Abstract

AI-assisted development has dramatically increased implementation speed, but not correctness. Rupify addresses this gap by turning requirements into executable, structured specifications that can be directly used by AI systems. Rather than relying on informal descriptions or heavyweight formal methods, Rupify operationalizes specifications as artifacts that can be generated, validated, and continuously enforced throughout development. Rupify is open source and available on GitHub: https://github.com/peterbb148/rupify

Why the name Rupify (RUP, UML, UCP)

Rupify takes its name from the Rational Unified Process (RUP), a structured approach to software engineering that emphasizes well-defined artifacts, traceability, and model-driven development. RUP uses the Unified Modeling Language (UML) to describe systems precisely through use cases, domain models, interaction diagrams, state machines, and deployment views. On top of this, Use Case Points (UCP) provide a way to estimate system size and effort based on functional structure rather than code.

Rupify operationalizes this chain—RUP for structure, UML for representation, and UCP for measurement—by turning it into an executable pipeline. Instead of producing documentation, it produces machine-interpretable models that AI systems can use directly for generation, validation, and estimation.

The Problem

AI systems are highly effective at generating, refining, and reviewing code, but they still depend on incomplete requirements, ambiguous intent, and inconsistent structure. This creates a fundamental mismatch where high-capability implementation systems operate on low-fidelity input.

The consequences are predictable. There is drift between intent and implementation, outputs vary across iterations, and correctness cannot be verified in a systematic way. Speed increases, but confidence does not.

The Idea Behind Rupify

Rupify introduces a structured, executable middle layer between intent and implementation. The process moves from interview to structured model, from model to executable artifacts, and from there into implementation and continuous validation.

The core idea is simple but fundamental. Specifications are not written primarily for humans; they are compiled for machines. Instead of acting as passive documentation, they become active inputs to the system.

What Rupify Does

Rupify provides a deterministic pipeline that starts with understanding a problem and ends with verifiable artifacts. Requirements are captured through structured interviews and translated into a canonical project model. From this model, Rupify generates RUP-aligned artifacts such as use cases, domain models, interaction diagrams, state models, and deployment views.

These artifacts are not static descriptions. They form the basis for use case point estimation and enable continuous validation against the original intent. The output is not just text, but a model that can be executed, tested, and checked.

Positioning

Rupify sits in the space between informal and formal approaches. On one side are notes, tickets, and lightweight specification formats. On the other are formal methods such as Z, TLA+, Alloy, and RAISE.

It provides structure without requiring full formalization, making it practical for real-world teams that need both speed and rigor. It is designed for environments where AI is already part of the workflow, but where correctness still matters.

Why This Matters Now

AI has shifted the bottleneck in software development. Writing code is no longer the primary constraint; defining correctness is. Without a structured specification layer, AI amplifies ambiguity rather than resolving it. Increased speed leads to increased drift, and verification becomes reactive instead of proactive.

Rupify addresses this by making correctness part of the input rather than an afterthought.

From Specification to Execution

Rupify enables a direct path from specification to execution. The generated artifacts are testable, traceable, and reproducible. Requirements can be followed through to implementation, estimates can be derived consistently using use case points, and systems can be continuously checked for conformance.

This allows AI agents to operate within clearly defined constraints instead of improvising from loosely defined prompts.

Practical Workflow

A typical workflow begins with a structured interview to capture intent. This is transformed into a canonical model, which in turn produces RUP artifacts. From these, estimation is derived and implementation is guided or generated. Throughout the process, validation is continuous and tied back to the specification.

The important shift is that every step is machine-interpretable and part of a coherent system.

Beyond Documentation

Traditional specifications are written, read, and eventually become outdated. Rupify specifications are generated, executed, and remain active parts of the system. They do not sit beside the implementation; they shape and constrain it.

Outlook

Rupify represents an early step toward a broader shift in software engineering. It points toward specification-driven development, where AI systems operate within executable intent and validation is built into the workflow.

The long-term direction is a move away from code-first development toward systems where specifications define, generate, and continuously validate the implementation.

« Older posts

© 2026 Peter Birkholm-Buch

Theme by Anders Noren — Up ↑