Skip to content
A·SOLODKAIA

Blog

The Traditional Architect's View on Determinism in AI Solutions

Artificial IntelligenceTechnology ArchitectureEnterprise Technology

Diagram showing inputs - people, data, and documents - flowing through an AI layer, then a verification step, into different types of outputs.

Why this discussion exists

AI is already being used to solve increasingly complex and serious business problems. However, everyone who used to work with Large Language Models for more than a few hours is aware of their limitations. They can produce different answers to the same question, misinterpret context, confidently return incorrect information, etc.

This is why scepticism around AI is understandable in enterprise world. Some of these problems are manageable for personal tasks, but they might cause serious financial harm for organizations. This is why one of the recurring questions around enterprise AI usage is whether its behaviour can be made predictable enough to trust it in production.

My recent conversation with the founder of an AI startup developing deterministic AI solutions gave me an idea to formalize my thoughts about this topic from a standpoint of a traditional solution architect. There are a lot of topics surrounding AI usage in production environments, many of them are serious enough to require lengthy discussion - for example, AI security. However, this article focuses on a particular aspect of AI systems: nondeterminism.

What is important for me as an architect is not only how determinism can be introduced into AI-based solutions, but also what we gain from it and what we sacrifice in return. And - most importantly - is determinism by itself the right thing to optimize for?

Where determinism can be enforced

Determinism of an AI solution means that it provides the same answer every time you enter the same input.

Broadly, there are several standard ways to make an AI-based solutions more predictable:

  • AI is not used during execution. Instead, it is used for producing software. Nondeterminism appears only at build time.
  • AI is being used locally for small, confined tasks such as parsing documents, the conventional software controls the whole process. Nondeterminism appears only at small steps such as data extraction.
  • AI translates human intent into predefined software operations. Choosing the action and its parameters introduces nondeterminism.
  • AI is used for complex tasks and produces broader results, but traditional software verifies the result before accepting it. The whole process is nondeterministic, only the acceptance test is deterministic.

Many of the existing solutions on the market are using those approaches - in many cases in a hybrid way.

Deterministic does not mean correct

Receiving consistent results from the same input already sounds like a great promise. Even if you set the temperature to zero for existing LLMs, it increases the probability of receiving the same answer repeatedly. But increased probability does not mean certainty. In practice, the provider processes your request together with other users' requests, and the arithmetic gives slightly different results depending on how many requests are in the batch, which depends on server load at that moment. This is why guaranteed determinism can be reassuring.

However, determinism is a weaker guarantee than it sounds. Firstly, if a deterministic system returns an incorrect answer, it returns this incorrect answer every time.

Secondly, determinism guarantees the same output only for identical input. Real customers never phrase things identically: "I lost my card", "I can't find my card" and "my card is gone" are the same request for the business. But they are three different inputs for the model and can resolve in different actions. A fully deterministic model can still handle them differently.

Lastly, there is a time factor plays key role: the models change and replace each over time, and inevitable this changes the behaviour and output.

Therefore, determinism alone is not enough to evaluate an AI-based architecture.

Are there any other important qualities to consider?

When we are speaking about introducing an AI solution for solving real world business problems, there are several characteristics we would typically want to look at.

The first four describe the quality of the solution:

  • Correctness
  • Auditability
  • Flexibility
  • Change velocity

And the other two describe economic-related qualities:

  • Development cost
  • Runtime cost

This is not an exhaustive list: for specific problems we might need to introduce additional ways to assess the same solution. For example, in many cases you would want to look at something like coverage - the percentage of cases which the solution is able to cover without human involvement.

Let's look at each of these qualities - their context and whether they are important to you or not.

Correctness

The ability of the system to generate the correct result.

This quality matters a lot when financial stakes are high - for example, in banking or anything which includes money, security and regulation. In contrast, for brainstorming and productivity applications correctness might be less important.

The priority indication for an architect: if an incorrect result could cause major financial damage, regulatory exposure, customer harm, or produce an irreversible action.

Auditability

The auditability consists of two parts: reconstruction how steps were executed and explanation how a particular result or decision was produced. Asking the model to explain itself produces a plausible story, not the actual cause. In practice, an AI step cannot be reproduced by re-running it, so its actual output has to be stored at the moment of the decision.

This quality is especially important for regulated areas of a business such as legal, healthcare, government, banking and others. It is less important for personal productivity, ideation and non-binding recommendations.

The priority indication for an architect: if someone may later ask, "why did this system make this decision?" and organization is obliged to answer.

Flexibility

How well the system handles inputs, situations, and variations that were not explicitly designed in advance.

The knowledge work, support systems, and research are among those areas which require this quality. In contrast, it might be less important for well-defined processes.

The priority indication for an architect: if it is difficult or unrealistic to enumerate most real-life scenarios before building the system.

Change velocity

How quickly the behaviour of the system can be changed when business rules, products, policies, or external conditions change.

It is important in businesses which need to respond rapidly to market changes: e-commerce, pricing, marketing, fraud response, product operations, etc. It is less important for mature back-office processes, stable accounting or reconciliation flows, and systems in which business rules rarely change.

The priority indication for an architect: the business expects behaviour changes weekly or monthly, and a normal development and release cycle would become a bottleneck.

Development cost

The cost of designing, integrating, testing and validating the solution before it starts delivering value.

This figure is critical for startups where the hypothetical cost benefit is yet to be proven. Experimental automation, low-volume workflows and internal tooling projects should take a close look at this property too. But it has less priority for high-value core processes where the cost benefit is calculable.

The priority indication for an architect: the hypothetical cost benefit is yet to be proven, or the cost of building this solution is very close to the realistic value it may create.

Runtime cost

The cost of processing each real request or transaction after deployment.

This figure matters the most for high-volume products. Customer support, content processing, transaction monitoring are examples of such areas. In contrast, high-value and low-volume processes, such as complex B2B cases or specialist analysis might consider this figure to have a lower priority.

The priority indication for an architect: this is a simple mathematics: even a small per-request cost may drain the budget when multiplied by millions of transactions.

Shifted responsibility

There is one additional phenomenon engineers and architects coming from traditional software may not expect.

We have been operating in the traditional software world long enough to have a fairly clear idea of how responsibility is distributed across a system. The boundaries are usually explicit, and when something goes wrong, it is normally possible to trace where the failure happened.

With AI-powered systems, it is much easier to overlook a shift in responsibility and create serious legal or operational problems when something goes wrong. A small interpretation error from an LLM can propagate into downstream systems and result in a much larger business impact. This reduces predictability and increases risk.

While technical responsibility shifts, the responsibility to customers and regulators does not change. The organization therefore has to define the internal side explicitly: who owns each AI step, what accuracy must be demonstrated before it goes live, what is monitored in production, and what the contract with the AI vendor says when the vendor's model is the source of the error. Standard terms usually leave that risk with the customer.

So, for every important property of the system, we should ask two additional questions:

  • What does this property cost us as an organization?
  • If it fails, which part of the system, - and which party - was responsible for preventing that failure?

Now let's look at common architectures

Many production-grade solutions combine several of the approaches listed below. It is therefore important to understand which approach is used for each business use case, rather than treating the whole system as a single architectural model.

Many of the tools we already use are built this way. For example, in Microsoft Copilot Studio the model decides which topic, tool or knowledge source to use, while the topics themselves are authored dialogue steps, and agent flows execute fixed, predefined logic.

Experienced users of existing LLMs also tend to adjust their expectations of reliability over time. Through repeated use, they learn where generated output is likely to be trustworthy and where it still needs verification. In enterprise systems, however, relying on this learned human judgement is not possible: the architecture itself must make those reliability boundaries explicit.

Now let's look at our building blocks and how we can assess their influence in our use cases. It is important to note that the ratings below compare the four approaches with each other, not with building the same system without AI.

Diagram summarising four architectural patterns for using AI in production: build then run, AI for interpretation, intent to actions, and propose then verify.

To use AI only to generate the workflow

How it works: AI generates a traditional software product.

What it gives you:

  • Correctness: potentially very high, provided the generated artefact and the underlying business rules are correct and properly tested.
  • Auditability: high – the software product artifact is fixed once built, and so the workflow can be audited and traced.
  • Flexibility: low. The production system can only support scenarios which were anticipated when the workflow was created.
  • Change velocity: relatively low but might be acceptable if business rules change rarely. On the other hand, it is not affected when the provider retires or updates a model.
  • Development cost: usually high. The cost is in formalizing the rules, for example collecting eligibility rules that are scattered across three legacy systems and several senior underwriters and writing test scenarios. The code generation part is cheap.
  • Runtime cost: very low, since the expensive model's calls are not a part of the workflow.

Major responsibility shift:

The responsibility for correctness and predictable behaviour shifts from AI system to the quality of specification, validation, software development and delivery discipline. Many organizations - for example, legacy banks - might find this too difficult or expensive if business rules are fragmented across legacy software, documentation and individual employees.

To use traditional software for the workflow, and AI for small, specific tasks

How it works: the whole solution is based on traditional software which utilizes AI for tasks like document parsing, or anything else small and well-defined.

What it gives you:

  • Correctness: potentially high, distributed between AI, validations and implemented business logic in traditional software.
  • Auditability: high, since the main workflow is implemented as traditional software, it can be made comprehensively traceable.
  • Flexibility: medium, permits more potentially unstructured inputs: for example, wording and formats, but the workflow is fixed.
  • Change velocity: medium. The AI part can be tuned using prompt or model configuration, but the process change would require traditional development.
  • Development cost: high. In practice it is traditional workflow development cost with additional cost for evaluating each AI step: test sets, accuracy measurements, resolution paths for answers below confidence threshold.
  • Runtime cost: medium: expensive model requests are required for execution, and documents processing is token heavy. A ten-page scanned bank statement costs many times more than a one-line customer message.

Major responsibility shift:

Responsibility is distributed across implementation layers. This is often one of the hardest architectures to reason about because no single layer guarantees end-to-end correctness on its own, and errors are silent. Since some LLM interpretation mistakes can propagate into downstream deterministic software, failures can become harder to predict and analyse.

Consider the simple example. If the model reads a salary of 21,000 as 12,000, the downstream code receives a perfectly valid number and declines the loan. The decline looks like a normal rule-based decision, so nobody investigates. The opposite error is worse: a salary of 12,000 read as 21,000 gets a loan approved that the customer cannot afford, and it surfaces only when the customer defaults. The main defence is cross-checking against an independent source, for example the salary on the certificate against the salary credits on the bank statement.

To use traditional software for the workflow, uses AI to understand human's intent

How it works: use AI to understand what the user wishes to do, execute the traditional software to do the action.

What it gives you:

  • Correctness: medium-to-high: if the model chooses the right action or parameter, the action execution will be correct if implemented correctly in traditional way. But there is risk of choosing a wrong action or parameter, which in some products can be reduced by human-in-the-loop.
  • Auditability: high, the workflow is easy to log.
  • Flexibility: medium in accepted inputs, low in business actions since the latter are strictly limited to a finite list.
  • Change velocity: medium – new business action would require explicit development as part of relevant integrations.
  • Development cost: medium to high. If all existing APIs already exist – the cost can be pleasantly low. However, in the case of an old and unstructured legacy landscape, it can grow very high.
  • Runtime cost: low to medium: the AI is being used only for intent recognition.

Major responsibility shift:

Responsibility partially shifts from building new business logic to correctly mapping human intent onto existing business capabilities.

AI is responsible for choosing the right action and its parameters, while the underlying deterministic system is responsible for executing it correctly. For some systems - for example, customer support - sensitive actions such as opening a new account, blocking a card or changing account settings can benefit from an additional confirmation from the customer.

To let AI propose the solution, then verify it before execution

How it works: AI proposes a solution. The traditional system then passes it to a verifier (traditional software, a person or another AI), and only a verified result is executed.

What it gives you:

  • Correctness: from medium to high, depending on verifier chosen. In many cases the reliable verification is hard or close to impossible to implement.
  • Auditability: medium to high – AI proposal, verification, result and final decision can be reconstructed and traced, however explaining it is significantly harder. The verifier's decision is easy to explain when it is software, but for complex tasks including AI-based or human-in-the-loop verifiers, explanation constitutes an additional challenge.
  • Flexibility: high, since AI can solve a larger scope without restriction before running verifier.
  • Change velocity: high, if verifier stays the same – there is no need to implement every single use case through specification or explicit use-case development.
  • Development cost: medium to high – depending on verifier. If the verifier already exists, for example a comprehensive automated test suite that checks business outcomes, the development cost can be low. If a custom verifier is required, each change might be very costly – especially if there is no strict pre-defined algorithm for verification.
  • Runtime cost: medium to high – AI generation, verification and regeneration might consume significant amount of resources.

Major responsibility shift:

Responsibility shifts from trusting AI to be correct to trusting the verifier to correctly validate solution properties we care about. The latter can itself become a complex task. Some open-ended questions - for example, whether a proposed business strategy is correct - may be very difficult to verify objectively.

AI-assisted programming is a good illustration: code that compiles is only proven to be valid code, not correct code. A function that applies an annual interest rate as a monthly one will not produce any compilation errors.

AI-based validation can be used as a verifier, but it does not provide deterministic assurance. It may increase confidence in the result, but the validation stays probabilistic. The same applies to human reviewers. A person who reviews two hundred AI-prepared cases a day, and finds almost all of them correct, gradually stops checking and starts approving. A human in the loop is only a real verifier if they have the time, the information and the incentive to disagree.

Diagram showing where responsibility shifts for each of the four architectural patterns: to specification and delivery discipline, to AI interpretation and business logic, to intent mapping and execution, or to the verifier.

Examples

In this section I show how I apply a tradition architecture design methodology to AI-based systems. The actual implementation would require high degree of knowledge for surrounding systems, overall architecture and product requirements.

Example 1

A legacy bank wants to introduce an AI-based customer channel for a loan application, using existing eligibility and pricing rules.

Diagram of a legacy bank loan application: customer and application documents flow through an AI layer into the pricing, eligibility, and lending systems, with correctness, auditability, and reuse of existing deterministic capabilities called out as priorities.

Importance of qualities:

  • Correctness: critical - mistakes can have direct financial and regulatory impact.
  • Auditability: critical - lending decisions may need to be reviewed later. For example, in the EU, AI systems used to assess the creditworthiness of individuals are classified as high-risk under the AI Act (Annex III). Under the same act an AI system that only performs a narrow procedural or preparatory task, without materially influencing the decision, may fall outside the high-risk category.
  • Flexibility: useful for documents and customer input but should not compromise correctness.
  • Change velocity: important - pricing and lending rules may change quickly to catch up after market changes.
  • Development cost: less critical since the product already has predictable value.
  • Runtime cost: relatively less important because each application has high business value.

Architecture fit

  • Uses AI only to generate the workflow: risky if business rules are fragmented across legacy systems and people. Can slow down reaction to market changes.
  • Uses traditional software for the workflow, and AI for small, specific tasks: strong fit for document processing and unstructured input. But document interpretation errors can propagate and create significant risk if validation does not catch it.
  • Uses traditional software for the workflow, uses AI to understand human's intent: looks like the strongest fit if pricing, eligibility and loan systems already exist. In this option the AI collects the customer's data and fills in the application; the existing eligibility engine makes the decision.
  • Lets AI propose the solution, verifies it before execution: high risk: for lending decisions, building a verifier capable of independently validating the proposed result may itself be a difficult task.

Conclusion:

For a legacy bank, I would usually prioritise reusing existing deterministic capabilities and adding AI around them, rather than replacing them. This choice, however, does not provide good change velocity – this is a necessary trade-off that needs to be made.

Example 2

Build a customer-support assistant that can perform simple user actions, not only answer questions.

Diagram of a customer-support assistant: customer requests are mapped to actions in an existing system via intent mapping, and factual questions are answered from a trusted source rather than generated freely.

This is a natural case for mapping human intent to existing system actions. AI interprets different ways of expressing the same request and maps them to operations already implemented by the underlying system.

For example:

"Freeze my card" → FREEZE_CARD

The first risk is incorrect interpretation: the system may execute the selected action perfectly while the AI has selected the wrong one. Should "I lost my card" be resolved into FREEZE_CARD or REPORT_LOST action? Freezing is temporary and can be undone if the card turns up; reporting it lost cancels it permanently and issues a new one.

For sensitive actions such as blocking a card, opening an account or changing important settings, the customer can remain in the loop:

"'I will permanently cancel the card ending 1234 and send you a new one. Confirm?"

This does not remove the need for normal business validations, but it provides an additional check that the AI understood the customer correctly.

The second risk is related to questions that require factual answers. An AI mistake here can mislead the customer and potentially create financial or regulatory risk.

The same intent-to-action mapping can be applied here: instead of letting the LLM generate the answer itself, the customer's question is mapped to a predefined operation that retrieves the answer from a reliable source.

The AI no longer writes the answer. It only picks one of the approved ones, but it can still pick the wrong one, for example the domestic transfer fee for a question about sending money abroad. So, the answer should say what it refers to.

The trade-off is the additional integration work required to support the expected set of product-related questions.

Conclusion

AI solutions are increasingly positioned not simply as productivity tools, but as products that can solve real business problems with a high degree of assurance. In that context, determinism and predictability are attractive because they are easy to understand, easy to demonstrate, and easy to sell.

But they are only part of what the business needs.

A solution can be highly predictable and still be wrong, too rigid for the process, too slow to adapt, too expensive to operate, or impossible to justify to a regulator. The business outcome depends on the whole architecture, not on one property of the model.

In practice, many of these systems are not deterministic end to end. They provide bounded nondeterminism: probabilistic behaviour is allowed in specific parts of the system, while architecture limits how far it can influence business-critical decisions or actions.

So when evaluating such a solution, the important questions are not only how deterministic it is, but also:

  • Where is nondeterminism still present?
  • What limits its propagation?
  • What does correctness depend on?
  • Does it stay that way when the model is updated?
  • What trade-offs were made to achieve this level of assurance?
  • And how responsibility is split between the organisation adopting the solution and the vendor?

The goal is not necessarily full determinism. The goal is to build a system whose behaviour is predictable, correct, auditable and flexible enough for the specific business problem - at a cost the organisation is prepared to accept.

← All blog posts