Skip to content

Optimization & governance

Use this guide to review optimization candidates, evaluate their evidence, move an approved intervention through its lifecycle, and verify realized savings. Attribution is the prerequisite. Every recommendation is advisory by default, confidence-gated, auditable, and reversible; Venturi does not change a customer system without an authorized customer workflow.

Optimization recommendations

Venturi turns the attribution graph into concrete, defensible opportunities:

Opportunity What Venturi does
Model right-sizing Identifies over-provisioned workloads and surfaces a cheaper model candidate, with workload-specific evidence from your own shadow data and required validation before any change.
Price-reversal avoidance Finds cases where token expansion or reasoning overhead makes a lower-list-price model more expensive in practice.
Idle or redundant usage Detects duplicate inference paths, idle deployments, and spend with no owner.
Provider or tier mix Compares equivalent provider, region, and service-tier options within the workload’s constraints.
Commitment opportunities Identifies stable usage that may qualify for a lower effective rate under a provider commitment.

No recommendation without workload evidence and your validation

Venturi never recommends a model swap on the promise of savings alone. A cheaper model is only surfaced as a candidate with workload-specific evidence from shadow data, and requires your validation before any change. Language is conservative throughout: you see an “optimization opportunity,” never a “savings guarantee.”

Each recommendation carries a rationale, an estimated saving, and the full interpretation metadata: confidence, evidence basis, and the data it rests on, so an engineer can evaluate it before acting.

Savings eligibility

Not every opportunity rests on the same strength of evidence, so every recommendation discloses a savings-eligibility state computed from the same confidence object that gates chargeback. You always know which bucket a recommendation sits in before you act on it:

State Confidence What it means for you
Billable coper ≥ 0.80 The underlying attribution is chargeback-grade. If you apply this swap, the realized savings can enter the verified-savings base.
Advisory only 0.50–0.79 (the twilight band) A real, actionable opportunity, but the attribution isn’t yet chargeback-grade. Apply it if you like; it just won’t be counted toward a success fee until the attribution qualifies.
Low confidence coper < 0.50, or less than 7 days of observation Surfaced for awareness; not a basis for savings claims yet.

The twilight band is honest, not hidden

Recommendations in the 0.50–0.80 band can look just as confident as chargeback-grade ones. Venturi labels them explicitly so you can tell which recommendations rest on chargeback-grade attribution and which don’t. An advisory-only recommendation that later improves to coper ≥ 0.80 is automatically re-qualified for the verified-savings base; you don’t lose the opportunity, it just becomes billable once the evidence catches up.

Advisory by default, active by choice

Optimization follows a deliberate, opt-in maturity progression. No mode change ever blocks production AI traffic.

Passive
observe & attribute

Advisory
recommend with rationale

Active
routing, opt-in

  • Passive: Venturi observes and attributes. Nothing is recommended or changed.
  • Advisory: Venturi recommends optimizations with rationale and verified equivalence. You decide what to apply.
  • Active: for opted-in workloads, Venturi can route to a verified-equivalent model. This is explicit, reversible, and never the default.

Each transition is opt-in, and the interceptor remains fail-open at every stage; moving to a more active mode can never cause an AI request to be blocked.

The intervention lifecycle

Every recommended or applied action (an “intervention”) moves through an explicit, auditable lifecycle with human control at each step:

Text Only
PROPOSED → APPROVED → ACTIVE → COMPLETED
                 └─→ REJECTED
  • A recommendation starts as PROPOSED.
  • A human with the right role approves or rejects it.
  • High-impact actions (for example, enabling active routing or a large budget change) require N-eyes approval and respect separation of duties: the person who proposes a change cannot be its sole approver.
  • Approved actions become ACTIVE, then COMPLETED, with the full history preserved in the audit log.

Human-controlled by construction

No AI subsystem in Venturi ever makes an autonomous change to your systems. Optimization recommendations are explainable, confidence-gated, human-reviewable, auditable, and reversible. Categories that are prohibited from automation are blocked by construction, not by a setting you could accidentally flip.

Savings evidence

A projected saving is a hypothesis. After an intervention, Venturi calculates realized savings against the reconciled bill and records the calculation in a versioned receipt that can be reviewed independently.

For every applied intervention, Venturi computes realized savings as the usage-normalized counterfactual cost (what the workload would have cost under the old model) minus the reconciled actual cost after the change, over each true-up period. The result is a deterministic, versioned savings-realization receipt (the savings analogue of your chargeback receipt) rendered in the evidence drawer and linked directly from the success-fee line on your bill.

What the savings-realization receipt shows you

  • The frozen baseline window and who froze it, so the counterfactual can’t drift under you.
  • The usage-normalization basis: savings are measured per workload volume, so organic growth in usage is never counted as savings.
  • Per-period counterfactual and reconciled-actual figures, and the realized-savings delta between them.
  • The share of underlying spend that is chargeback-grade (coper ≥ 0.80): the only spend that counts toward the savings base.
  • The equivalence assertion behind the swap, and the methodology version, pinned to a specific model version.

Realized savings

The receipt is the sole input to any savings-share billing. Venturi never bills on a projection or a dashboard assertion:

  • Spend below the coper ≥ 0.80 chargeback floor is excluded from the savings base, with the reason shown on the receipt. A success fee never rests on attribution that isn’t chargeback-grade.
  • If the counterfactual or the volume normalization can’t be resolved, the line is held from billing rather than billed on a silent default: an honest unknown, never a guess.

Degraded attribution

If Venturi’s decision-time interceptor ever falls back to its fail-open path on the AI hot path, those attributions are marked degraded, and degraded spend is excluded from both the frozen counterfactual and the post-change actual. The receipt reports the share of degraded spend it had to exclude. If that share crosses a configurable threshold (10% by default), the receipt is marked provisional and is not billable until the window is re-observed under normal capture.

A degradation episode can’t masquerade as a saving

A success fee billed partly on spend the decision-time interceptor fell back on would be the most disputable number Venturi could produce. By excluding degraded attribution outright, a realized-savings figure can never be an artifact of a fallback episode that happened to overlap your post-change window.

Workload-scoped equivalence

The equivalence behind a billable swap isn’t “it’s on a list.” Every equivalence assertion that enters the verified-savings base carries the specific task type and workload scope it’s asserted for, an assertion date and source version, a strength/confidence, and (for empirical shadow evaluations) the sample and the metric deltas. The receipt embeds the exact assertion it relied on.

If a swap’s equivalence assertion is out of scope for your workload’s task type, or has gone stale beyond a configurable freshness window, the recommendation is automatically downgraded to advisory-only and removed from the billing base until it’s re-verified. When you dispute a success fee, the receipt has something workload-specific to point to, not a generic claim of “verified.”

Energy and carbon impact

Where a swap carries an energy or carbon differential, it is computed only over the chargeback-grade (coper ≥ 0.80) workload it actually applies to, and labeled with whether the evidence is shadow or production. When the swap is applied, the realized carbon delta is reconciled into the same savings-realization receipt using your actual post-change energy profile, never the projected one. Where a model isn’t catalogued for energy, the carbon delta is shown as unknown rather than silently treated as zero, so the energy-aware leg holds up under procurement challenge.

Savings disputes

A savings-realization receipt (or the success-fee line it backs) is disputable through the same workflow as an attribution. The dispute doesn’t mutate the receipt; it resolves to one of three outcomes:

  • a corrected counterfactual or normalization, re-billed with an append-only audit trail and a re-verification that you’re never double-charged;
  • accept the figure as billed; or
  • close as duplicate.

The savings number is tied to a receipt

The success-fee model bills on verified realized savings against a frozen counterfactual, and every step is addressable, exportable, and contestable. The first contested invoice resolves through a structured dispute path, not a manual escalation, because the number was audit-grade from the start.

For the confidence floors, retention, and capture guarantees these receipts inherit from attribution, see Cost attribution & chargeback and Trust & security.

Budget governance

Govern AI spend with per-team and per-project budgets:

Active (Gate) sends a signed deny recommendation to a customer-controlled enforcement point; it never converts Venturi’s fail-open interceptor into a blocking proxy.

  • Advisory by default. Budgets alert when consumption crosses configured thresholds, without blocking anything.
  • Customer-enforced stop by opt-in. Venturi can emit a signed deny recommendation for a named workload. A customer-controlled enforcement point may apply it; Venturi’s fail-open interceptor never blocks traffic.
  • Threshold alerts fire at 75% / 100% / 150% of plan, delivered to email, Slack, Microsoft Teams, or a signed webhook.

Budget configuration, breach handling, forecasting, and alert routing are covered in depth in Budgets & alerts.

Governance you can demonstrate

For any single attribution or recommendation, Venturi can show a CISO, a FinOps lead, an auditor, or a regulator exactly what the AI concluded, with what confidence, on what evidence, with which model version, who could override it, and where the human-control boundary sits, reachable in a couple of clicks from the result. The governance console surfaces the evidence, confidence, approval state, and realization status of each Venturi recommendation. Disputes and overrides are first-class, audited workflows. Provider model-health or drift data appears only when a configured source supplies it; Venturi does not infer an authoritative deployment inventory from billing records alone.

Venturi’s published scope and non-goals match what the system actually enforces, so the boundary you’re shown is the boundary that holds.