Skip to main content
Choose AI Stack
Search

Positioning guide · Updated 2026-09-13

AI software stack vs AI infrastructure stack.

Use this guide to decide whether your next action is buying and governing software or assigning an engineering-owned platform decision. The boundary changes who approves the cost, who operates the system, and what must be resolved before production.

Decision areaAI software / workflow stackAI infrastructure / MLOps stack

Primary buyer question

Which AI tools should a person or team use for a workflow, and what should they avoid for now?

How should an organization build, host, monitor, evaluate, and operate AI systems?

Typical examples

Chat assistants, coding assistants, meeting tools, research tools, workspace AI, and workflow-specific combinations.

Model hosting, vector databases, data pipelines, evaluation systems, observability, orchestration, and MLOps platforms.

Cost to approve

Seats, subscriptions, usage allowances, implementation time, training, and the cost of overlapping tools.

Compute, storage, model usage, data movement, observability, engineering capacity, reliability work, and on-call ownership.

Operating owner

A workflow or business owner can run the pilot when the vendor operates the service and internal work is mainly adoption, permissions, and administration.

An engineering or platform owner is required when the team must operate hosting, data pipelines, evaluation, monitoring, scaling, or recovery.

Decision inputs

Role, workflow, team size, budget, privacy/security requirements, rollout risk, and alternatives.

Latency, scale, model selection, data architecture, compliance controls, deployment topology, and engineering operations.

Choose AI Stack coverage

In scope: practical AI software stack recommendations for tech workers and small teams.

Out of scope: model infrastructure architecture, MLOps platform design, and production AI system operations.

Mixed decision

Separate the pilot from the production obligation.

A software pilot can answer whether a workflow improves. It does not automatically approve the platform cost or operating model required for production. Record both decisions before the pilot is treated as a rollout commitment.

Who operates the system?

Keep the decision in the software lane when the vendor operates the finished service and your team mainly owns workflow adoption and administration. A managed model or API can still create platform work: if your team must build and own integration, evaluation, fallback, monitoring, or production reliability around it, bring in an infrastructure owner.

Which cost is being approved?

A software budget should include seats, usage, rollout, and duplicate subscriptions. An infrastructure budget must also include compute, storage, engineering time, observability, incident response, and ongoing model operations.

What blocks production?

A bounded software pilot may continue while a platform dependency is investigated, but production approval must pause when data flow, hosting, monitoring, recovery, or ownership remains unresolved.

Write down the handoff before approval.

Software decision
Record the workflow, shortlisted tool, decision owner, budget owner, and the result the pilot must prove.
Platform dependency
Name the unresolved hosting, data, evaluation, monitoring, reliability, or operating-cost question and assign its owner.
Release gate
State which dependency, cost estimate, and operating owner must be resolved before the software pilot can move into production, and record when pilot-only approval expires if that dependency is still open.

Make the next approval decision explicit.

Continue the software pilot

Proceed when the tool can be tested safely inside the current workflow and the unresolved platform question does not affect the pilot result.

Expire pilot-only approval

Set a reapproval date when infrastructure work is still open. If that date arrives without a cleared release gate, stop treating the pilot as approved continuation. To renew it, record the current dependency state, what changed since the last approval, the accountable owner, the next release gate, and why the pilot remains safe at the same bounded scope; otherwise pause it.

Reapprove after the dependency changes

Treat a material change to the unresolved platform dependency—such as data flow, hosting, monitoring, reliability, operating cost, or accountable owner—as a new approval boundary. Recheck whether the bounded pilot is still safe instead of waiting for the existing reapproval date.

Prove production clearance

Before production approval, record the resolved dependency, the validated operating-cost estimate, the accountable operating owner, monitoring and recovery readiness, and release-gate signoff. Prove recovery readiness on one representative failure or rollback path: name who can stop the rollout, what signal triggers the rollback, and how the workflow returns to the last safe pilot state. If that evidence is missing or any other clearance item remains unresolved, keep the work pilot-only.

Limit clearance to the approved production scope

Tie production clearance to the exact workflow, data boundary, deployment path, operating owner, and recovery obligation that were reviewed. Moving into a new workflow, team, data class, hosting path, or materially broader operating scope is a new production-clearance decision; do not inherit the old approval automatically.

Reaccept clearance after an owner handoff

When accountable operating ownership changes, require the incoming owner to explicitly accept the reviewed production scope, monitoring and recovery obligations, rollback authority, and release conditions before the prior clearance can remain in force. If the incoming owner cannot accept those obligations without changing the boundary, return the workflow to pilot-only and issue fresh production clearance.

Retire superseded production clearance

When fresh production clearance replaces an earlier record, mark the earlier clearance as superseded, record the successor clearance reference, and record the reason or trigger that required replacement—such as a changed dependency, cost, owner, monitoring or recovery posture, release condition, or production scope. Do not leave two records looking simultaneously valid: operators should be able to tell which clearance controls the current production scope, which approval was replaced, and what changed at the replacement boundary.

Revalidate stale production clearance

Give the production-clearance record a review date or trigger. Recheck it when that date arrives or when the dependency state, validated operating-cost estimate, accountable owner, monitoring or recovery posture, release conditions, or approved production scope materially change. If the prior evidence no longer matches the production obligation, return to pilot-only until clearance is renewed.

Route the platform decision

Send the dependency to the infrastructure owner with the pilot evidence, required production condition, budget question, and review date.