DECISION TOOLKIT · PLATFORM SELECTION

Position the tool on the work, not on the demo

The map is a starting hypothesis, not a benchmark. It separates speed-to-value from control and reversibility, then forces the team to validate the position on representative work.

01Illustrative positioning · validate with the same evidence pack
02Fit by workflow archetype

A single winner is usually the wrong output. Select a primary pattern, then define where a second pattern adds material value.

ArchetypeWorkspaceCoding agentDesktop agentOpen stackOrchestration
Broad knowledge workstrongstrongmediummediummedium
Software deliverymediummediumstrongstronglow
Desktop / local workmediumlowlowlowstrong
Differentiated productlowlowmediumstrongstrong

The evaluation unit is a completed workflow, not a model response.

“Which AI tool should we choose?” sounds like a product question. In an enterprise, it is an architecture decision involving data, identity, workflows, controls, economics and the ability to change direction.

Claude and ChatGPT Enterprise address broad knowledge work; Claude Code and Codex reshape software delivery; Kimi Work brings agent workflows to local files, browsers and recurring work. Kletron — the agent and multi-agent product built by Numezis and already used by SMEs — adds a complete orchestration layer for complex business missions. Open-weight architectures create still other control and integration options. These systems are not interchangeable, and a feature comparison alone is inadequate: the right answer depends on the workflow the organization is trying to operate.

Numezis works hands-on with these products across configuration, coding workflows, tools and MCP integrations, evaluation, permission design and production controls. Product fluency matters, but it must remain subordinate to the architecture decision.

Define the work before the product

Specify the workflow, users, systems touched, data categories, actions permitted and required outcome. Distinguish individual productivity, collaborative knowledge work, software engineering, customer-facing applications and autonomous system actions. Each creates a different trust boundary.

Do not ask which model is “best” in the abstract. Build a representative evaluation set from the actual work and define acceptance thresholds for quality, latency, reliability and cost.

Evaluate six decision dimensions

1. Data boundaries

Map inputs, context, outputs, telemetry and retained content. Establish where each travels, how long it remains and whether it may be used for service improvement. Contractual and technical boundaries must agree.

2. Integration

An isolated assistant may create local productivity but little operating leverage. Assess identity, permissions, APIs, MCP connectivity, retrieval, event systems and the ability to embed control into existing workflows.

3. Security and governance

Consider administration, access control, audit evidence, data-loss prevention, tool permissions, supply-chain dependencies, incident response and model change management. For agents, every connected tool expands the attack and failure surface.

4. Performance

Measure on the real task. A product may be excellent at drafting and weaker at coding, long-context analysis, structured extraction or tool use. Model quality also changes; evaluations must be repeatable rather than a one-off selection exercise.

5. Total economics

License cost is only one component. Include inference, integration, data preparation, evaluation, operational support, change management and the human review still required. Compare cost per successful workflow outcome, not cost per token alone.

6. Reversibility

Identify what becomes proprietary: prompts, agents, integrations, evaluation data, user behavior and knowledge indexes. Standards and abstraction can help, but unnecessary abstraction also creates cost. Preserve flexibility where the future option has material value.

Match deployment to risk and advantage

SaaS can offer the fastest adoption and strongest integrated experience. Cloud APIs provide greater product control. Local or self-hosted models can change data and operational boundaries, but transfer more responsibility to the enterprise. Hybrid architectures are often rational when different workflows carry different requirements.

The decision should be explicit:

PatternStrong fitPrimary trade-off
Enterprise SaaSbroad knowledge-work adoptionproduct boundary and vendor dependence
Managed model APIdifferentiated applicationsengineering and operations ownership
Self-hosted/open weightspecific control or economics needsfull lifecycle responsibility
Hybridvaried workflows and risk tiersarchitecture and governance complexity

Record the decision, including the exit

A useful architecture decision record states the context, options considered, evidence, chosen pattern, risks accepted and conditions that would trigger review. It also describes the exit path: how data, integrations and evaluations could move if the product, terms or performance change.

The goal is not permanent neutrality. The enterprise should make a strong current choice without accidentally making an irreversible one.

Start from the trust boundary

The most useful tool comparison begins by drawing the trust boundary around the work. What may the system read? What may it remember? What may it call? What may it change? Where must a person confirm? Which evidence must remain available after the workflow is complete?

This changes the shortlist immediately. A broad enterprise workspace may be the right adoption layer for drafting, analysis and knowledge work, while a coding agent belongs in an isolated repository workflow and a desktop agent belongs in a tightly scoped local-file experiment. An open architecture may be justified for a differentiated product, not because “open” is inherently safer but because the organization needs to own a control plane that a SaaS product cannot expose.

The right unit of comparison is a workflow pattern:

PatternTypical authorityPrimary proof obligationReasonable default
Assistread and draftquality and data boundarymanaged enterprise workspace
Retrieve & decideread, compare, recommendgroundedness, permissions, reviewer tracemanaged workspace or API
Engineerinspect repo, edit, test, proposesandboxing, diff review, test evidencecoding agent with isolated execution
Operatecall tools, update systemsleast privilege, validation, rollbackorchestrated application
Productiseserve external users at scalelifecycle, observability, unit economics, exitAPI or owned control plane

Build one evidence pack for every candidate

Procurement demos optimize for fluency. An evidence pack optimizes for decision quality. Use the same tasks, inputs, permissions, evaluation rubric and cost assumptions for every candidate. Include cases that expose the boundary conditions: incomplete source material, conflicting instructions, multilingual content, sensitive fields, prompt injection, tool failure and a request to take an irreversible action.

The pack should produce six artifacts:

  1. Task quality — correctness, completeness, groundedness, style and reviewer effort.
  2. Reliability — repeatability, timeout behavior, tool-call success and graceful failure.
  3. Control behavior — what the system refuses, what it asks to confirm and what it logs.
  4. Integration effort — identity, data connectors, APIs, MCP servers, environments and support burden.
  5. Economics — license or inference, implementation, supervision, operations and exit cost.
  6. Reversibility — what can be exported, what must be rebuilt and what becomes a proprietary dependency.

Do not collapse the pack into a single model score. A candidate can win quality and lose the decision because it cannot be integrated within the required data boundary. Conversely, a lower-scoring model can be the better enterprise choice when its controls, cost and reversibility create a stronger completed workflow.

Compare architectures, not just products

There are four recurring patterns.

Managed workspace. Strong default for broad internal adoption. It concentrates identity, administration, user experience and model access in one product. The trade-off is that the enterprise accepts the vendor’s product boundary and must examine how data, apps, tools, retention and model changes are governed.

Managed API. Strong default for a differentiated application or controlled workflow. The enterprise owns more of the user experience, orchestration, retrieval and policy layer. The trade-off is engineering and operational ownership: a model API does not provide a complete service by itself.

Coding or desktop agent. Strong default for work that genuinely needs local files, repositories, terminals or browser interaction. These agents can reduce friction because they operate where the work already lives. The trade-off is a larger action surface and more responsibility for sandboxing, local permissions, secrets, browser sessions and human review.

Open or self-hosted architecture. Strong default when control, residency, cost structure or product differentiation justifies full lifecycle ownership. The trade-off is not only GPU operations. It includes model evaluation, patching, capacity planning, abuse handling, security response, change management and a support capability that SaaS would otherwise provide.

Hybrid is often the rational answer, but only when the boundaries are explicit. Use the workspace for low-risk adoption, the coding agent for repository work, an API for a customer-facing workflow and an open model only where the control requirement is material. A random collection of tools creates duplicated identity, inconsistent data classes and no clear owner.

Make economics comparable at workflow level

The cost decision is rarely “seat price versus token price.” Model the full cost per successful outcome:

total workflow cost = platform + inference + integration + data preparation + evaluation + supervision + operations + change + exit

Then compare it with the current cost of completing the workflow, including rework and the cost of delay. A tool that is inexpensive per request can be expensive per completed case if users must repair most outputs. A premium platform can be economically sound when it removes integration effort and creates faster, safer adoption across many teams.

Run three scenarios: pilot, expected scale and stress. Stress should include higher volume, longer context, lower success rate, model migration and additional human review. The purpose is not to predict a precise number. It is to identify which assumptions could reverse the decision.

Reversibility is an investment choice

Reversibility is not free. Abstraction layers, portable prompts, shared evaluation sets and provider-neutral interfaces cost time. They are worth buying when a future change is plausible and expensive. They are not worth buying when they obscure the product capabilities that create today’s value.

Record the exit in five lines: data export format, integration inventory, evaluation portability, user and training dependencies, and the estimated time to a credible replacement. Review it when contracts, pricing, model behavior or business criticality changes.

The professional answer is not “we are vendor neutral.” It is “we know why this pattern wins for this workflow, we know what we are accepting, and we have kept the options that are worth keeping.”