Tactical Edge
Contact Us
Back to Insights

Data Contracts for AI: Catching Model Risk When Schemas Change

A renamed column or new enum value can degrade AI behavior before any data-quality check fires. Here is how to make data contracts enforceable with Glue, Iceberg, and CI gates.

Data & Analytics16 min
By David Chen, Principal Engineer ยท September 30, 2026
Data ContractsSchema EvolutionApache IcebergAWS GlueMLOps

A producer team renames status_cd to order_state and, in the same release, adds a new value called pending_review. Freshness monitors stay green. Row counts land within the usual band. Null rates do not budge. Nothing in the pipeline fails. Two days later, a support agent starts quietly routing a class of orders to the wrong queue because its routing logic branches on a set of four enum values and now sees five.

A data contract for AI is a versioned, machine-checkable agreement between a data producer and an AI consumer that covers schema, field semantics, enum domains, freshness expectations, and the process for changing any of them. Unlike a data-quality rule, it is owned by the producer, declared before the change, and enforced where a merge can fail. That last part is the whole point: a contract term that cannot break a build is documentation.

AI consumers are more fragile than dashboards because they encode field meaning in places no validator reads. A prompt says "the risk_band column is A through E, where A is lowest risk." A feature transform assumes amount is in cents. A retrieval filter assumes region uses ISO codes. A tool schema declares an enum inline. None of that lives in the schema registry, and none of it fails loudly when the upstream meaning shifts. Contracts are one control in a set, not a replacement for evaluation suites or trace-based observability. This article shows where each one belongs.

What a Contract Must Cover That Your Schema Registry Does Not

Split contract terms into three layers and the coverage gaps become obvious.

Structural terms cover field presence, type, and nullability. This is the layer registries handle well, and the layer most teams mistake for the whole problem.

Semantic terms cover meaning: unit, sign convention, enum domain, timezone, business definition, and the identity of the join key. A currency field can change from cents to dollars with no type change. A timestamp can change from event time to ingest time with no type change. An enum can gain a legitimate new value that is perfectly valid to the producer and catastrophic to a consumer that branches on values rather than passing them through.

Behavioral terms cover freshness SLO, backfill policy, late-arriving data windows, deletion semantics, and retention. An agent that reads a table refreshed hourly and one refreshed on demand behave differently even when every value is correct.

Contract layerWhat it declaresDetection mechanismSilent failure if missingOwner
StructuralField names, types, nullability, required setSchema registry compatibility check on producer mergeConsumer transform throws, or silently coerces a typeProducer platform team
Semantic (values)Enum domain, allowed ranges, units, sign conventionValue-domain rules in a data-quality ruleset, run per batchRouter or eligibility logic mishandles an unseen valueProducer domain owner
Semantic (meaning)Business definition text, join key identity, timestamp basisReview gate plus prompt/transform reference to the contract textPrompt describes a field that no longer means what it saysNamed field steward
BehavioralFreshness SLO, backfill policy, late-data windowPartition-age assertion in the consumer's pre-run checkModel answers confidently from a stale partitionProducer on-call
LineageDeclared consumers and their tierContract file registry, checked in CIChange ships with no notification path to the affected teamData platform owner

The mechanisms matter more than the categories. Use your schema registry's compatibility mode for the structural layer, a rule-based data-quality ruleset for value domains and distributions, and table snapshots for reconstructing the exact state a model read. AWS Glue Schema Registry and Glue Data Quality cover the first two on AWS; Great Expectations, Soda, and dbt tests cover the same ground if your stack is not Glue-centric. The tool choice is less important than whether the rule runs on a schedule you can prove.

One term that teams consistently forget: the declared consumer list. If a contract does not name who depends on it, a violation has no notification path, and the producer team finds out from a Slack message three days later.

Anchor the Analytical Contract to the Iceberg Spec, Not a Pipeline Convention

For years the analytical contract was a file-layout convention: partition by date, write Parquet, do not change column order, and hope. The Apache Iceberg table specification replaced that with a documented format. The spec defines table metadata, snapshots, and schema evolution, which is what makes time travel, concurrent writes, and engine independence possible [1].

The operational payoff for AI teams is specific. When a model behaves differently this week than last week, you do not argue from memory about what the table looked like. You read the snapshot the model actually consumed.

That turns one habit into a requirement: record the Iceberg snapshot identifier alongside every training run, feature materialization, and evaluation run. When a suspected contract violation surfaces, you can re-materialize the exact inputs, diff them against the current snapshot, and turn an anecdote into a reproducible test case. Without the snapshot ID, "the data changed" is an opinion.

Pair the spec with the governed-domain pattern from AWS data modernization guidance: publish one documented dataset per domain with named ownership, freshness expectations, and access controls, and make agents read only from it [7]. Agents that read from three different extracts of the same entity will disagree with each other, and no contract can reconcile that after the fact. This is the same discipline we apply in enterprise data and analytics platform work before any retrieval layer gets built.

One boundary worth stating clearly. Operational database migration decisions (rehost, replatform to a managed engine, or re-engineer per workload) are sequenced separately from analytical table modernization [7], and AWS DMS ongoing replication lets source and target run in parallel so reconciliation happens before cutover [11]. That separation is what lets you ship a contract for one domain this quarter while the wider estate migrates behind it. You do not need the migration finished to make one data path trustworthy. Our database modernization factory approach treats those as two tracks on purpose.

Making the Contract Executable: Producer CI, Consumer CI, Runtime

Three enforcement points, each catching a different class of failure.

  • Producer pull-request check: does the proposed change satisfy the compatibility the contract declares? Catches declared, intentional changes before they ship.
  • Consumer build check: does my transform, prompt, or tool schema still match the contract version I depend on? Catches drift in consumer code, including a prompt that describes an enum the contract no longer lists.
  • Runtime assertion: does today's batch or stream actually match the declared domain and freshness? Catches undeclared reality, like an upstream vendor feed that started emitting a sixth status value nobody announced.

Start with the contract itself. Keep it in the producer's repository, not a central wiki.

yaml
# contracts/orders/order_events.v2.yaml
dataset: analytics.orders.order_events
version: 2
owner: orders-platform@example.com
table_format: iceberg          # schema evolution per Iceberg spec
freshness_slo: PT1H
late_data_window: PT6H

fields:
  - name: order_state
    type: string
    steward: orders-platform@example.com
    semantics: "Current fulfillment state. Terminal states are shipped, cancelled."
    enum:
      closed: true             # new values are a BREAKING change
      values: [created, paid, shipped, cancelled]
  - name: amount_minor_units
    type: long
    semantics: "Order total in minor currency units (cents). Never negative."
    unit: minor_units
    min: 0

consumers:
  - name: support-triage-agent
    tier: 1
    depends_on: [order_state, amount_minor_units]
    branching_on: [order_state]   # consumer switches on values
  - name: revenue_dashboard
    tier: 3
    depends_on: [amount_minor_units]

The closed: true flag and the branching_on declaration are what make the check useful. Now the producer adds a value:

text
$ contractcheck diff contracts/orders/order_events.v2.yaml
FAIL  enum_domain_expanded  field=order_state
      declared: [created, paid, shipped, cancelled]
      proposed: [created, paid, shipped, cancelled, pending_review]
      enum.closed = true  ->  BREAKING under this contract

      Affected consumers (branching_on order_state):
        support-triage-agent   tier=1   owner=ai-platform@example.com

      Required before merge:
        1. bump contract to version 3
        2. acknowledgement from ai-platform@example.com
        3. re-run eval suite for support-triage-agent on v3

This is where the schema-registry compatibility mode alone falls short. A BACKWARD-compatible change is one that existing readers can still parse. Adding an enum value usually parses fine. It is only breaking for consumers that branch on the value, and the registry has no idea which consumers do that. The contract does, because the consumer declared it.

[1]
Snapshots specified
The Iceberg table spec defines metadata, snapshots, and schema evolution, so a model's exact input state is addressable rather than remembered
[2]
Attributes portable
OpenTelemetry generative AI semantic conventions give named span attributes for model and tool operations, so consumer-side signals survive a vendor change
[3]
Gaps enumerated
The NIST Generative AI Profile lists generative-specific risks and suggested actions, exposing what a predictive-model control set misses
[5]
Variance measured
tau-bench reports reliability across repeated trials of the same task, not a single pass, which is the method to reuse after a contract version change

The practical rule: no contract term is real until a merge can fail because of it. Write the term, then write the check that enforces it, in the same pull request.

Tiering Contracts by Blast Radius Instead of Table Popularity

The default instinct is to contract the most-queried tables first. That optimizes for the wrong risk. A table with forty dashboard consumers produces forty confused analysts who file tickets. A table with one consumer that drives an irreversible agent action produces a refund issued, a credit line adjusted, or a patient record routed incorrectly, with no human in the path to notice.

Tier by what happens downstream when the field is wrong.

TierExample dependencyRequired contract termsChange processGate on violation
1: Irreversible actionField drives an agent write: refund, order cancellation, entitlement changeClosed enums, units, freshness SLO, declared branching consumers, snapshot pinningVersion bump plus written consumer acknowledgement before mergeBlock the write path; agent degrades to read-only and escalates
2: Decision featurePricing, credit, eligibility, or risk scoring featureClosed enums, ranges, distribution bounds, backfill policyVersion bump plus 5-day notice to declared consumersHalt feature materialization; serve last known-good snapshot
3: Retrieval filterMetadata field used to scope a knowledge base queryEnum domain, join key identity, freshness SLOVersion bump plus notice; no acknowledgement requiredAlert on retrieval hit-rate shift; run degraded with wider scope
4: Explanatory contextField surfaced in an answer but not used for control flowSemantics text, nullabilityChangelog entryLog and notify; no blocking
5: Reporting onlyDashboard or scheduled exportStructural terms onlyChangelog entryLog

Tiering also maps cleanly onto governance obligations. The NIST AI Risk Management Framework organizes practice into Govern, Map, Measure, and Manage, which means measurement evidence and a defined response are expected outputs, not optional extras [4]. The Generative AI Profile is the fastest way to find the gaps a predictive-model control set misses [3]. If your organization operates in scope of Regulation (EU) 2024/1689, obligations attach per use case and per role, provider or deployer, which is one more reason the contract's consumer list needs to name use cases rather than services [9].

One scheduling change makes this work. Register review should trigger on new data sources and new tool permissions, not only on a quarterly calendar. A contract version increment is exactly that kind of trigger event, which is how data contracts and AI governance practice stop being two separate programs.

Tier Assignment Is a Consumer Decision
Producers almost always under-tier their own tables, because they see row counts and query volume, not downstream consequences. The consumer team that branches on an enum to decide whether to issue a refund is the only party that knows the blast radius. Make tier a field the consumer writes into the contract, require the producer to acknowledge it, and audit disagreements. A table nobody queries can still be Tier 1.

Detecting Semantic Drift Through the Consumer's Traces

Contracts catch declared changes. They do not catch a vendor feed that quietly starts emitting a new code, or an upstream system that reinterprets an existing value. For undeclared change you need the consumer's own telemetry, which means both controls are required and neither substitutes for the other.

Instrument retrieval steps, model calls, and tool invocations as spans using the OpenTelemetry generative AI semantic conventions so attribute names stay portable across vendors and backends [2]. Propagate one correlation identifier from the user request through every nested agent, subagent, and tool span so a full run reconstructs as a single trace [2]. If tools are exposed through a Model Context Protocol server, pin a specification revision and treat protocol upgrades as reviewed changes rather than dependency bumps [10].

Then watch for the signals that expose upstream data change on the AI side:

  • Retrieval hit-rate shift for a specific filter value. A region or status filter that used to return documents and now returns nothing is a renamed or re-coded field, not a model problem.
  • Tool calls with arguments outside a known enum. The agent is faithfully passing through a value the contract never declared.
  • Rising loop or repeat-call counts. Agents retry when a tool returns something that does not match what the prompt led them to expect.
  • Truncated or abandoned runs. Often the first visible symptom of a distribution change in a field the agent uses to scope its work.
  • Confidence or refusal rate moving on a stable scenario set. The scenario did not change, so something it reads did.

Close the loop. Archive full traces of failures and low-confidence runs, promote them into the offline evaluation suite, and re-run the suite whenever a contract version increments. Grade on machine-checkable end states rather than reference answer strings, following the outcome-verification pattern SWE-bench established by running tests after an agent's patch [6], and score policy adherence as its own dimension the way tau-bench does alongside task success [5].

The variance point deserves emphasis. tau-bench reports reliability across repeated trials of the same task rather than a single pass [5]. High run-to-run variance on one scenario is a concrete blocker for expanding an agent's write scope, and a contract version change is a reason to re-measure variance rather than assume it held. This is the same measurement loop described in our agentic AI approach, applied to the data layer instead of the prompt layer.

A 30-Day Sequence to Get One Contract Into Production

Do not start with a data-contract platform. Start with one use case and one contract that can fail a build.

Week 1: Trace backward from one AI use case. Pick the agent or retrieval path with the largest blast radius, not the most traffic. List the exact tables and fields it reads. Then read your own prompts, transforms, and tool schemas and write down every field semantic they already imply: units, enum sets, timestamp basis, join keys. This step usually surfaces two or three assumptions nobody wrote down.

Week 2: Publish version 1 in the producer's repository. Assign a named steward per field, not a team alias for the whole dataset. Register declared consumers with their tier and whether they branch on values. Get the producer to acknowledge the tier assignments in writing, including the ones they disagree with.

Week 3: Wire the enforcement points. Add the producer pull-request check. Add a value-domain and freshness ruleset that runs per batch. Record the Iceberg snapshot ID in every consumer run, materialization, and evaluation so the input state is addressable later [1]. Review the retrieval and serving path against the AWS Well-Architected Generative AI Lens while you are in the code [8].

Week 4: Prove the suite fails. Add three adversarial evaluation scenarios driven by contract violation: an unknown enum value, a unit change from minor to major currency units, and a stale partition outside the freshness SLO. Inject each violation and confirm the suite goes red. A contract check you have never seen fail is a contract check you do not have.

The metric to start tracking this week: the share of AI-consumed fields that have both a named steward and a declared contract version. Count it by field, not by table, and report it per tier. Most teams find the number is close to zero for Tier 1 fields and comfortably high for reporting tables, which is precisely the inversion this article is about.

Frequently Asked Questions About Data Contracts for AI

What is the difference between a data contract and a data-quality check?

A data-quality check tells you that today's data looks wrong. A data contract tells the producer, before they merge, that a change they are about to make will break a named consumer. One is detection after the fact; the other is a gate with an owner on both sides.

Do data contracts require a specific tool?

No, but the work divides cleanly across three mechanisms. A schema registry (AWS Glue Schema Registry, Confluent, or Apicurio) enforces structural compatibility when a producer registers a new schema version. A rule-based data-quality engine (Glue Data Quality, Soda, Great Expectations, or dbt tests) enforces value domains, ranges, and freshness per batch. The Apache Iceberg table specification provides the snapshots and schema evolution semantics that let you reconstruct the exact state a model read [1]. The contract file itself is just YAML in the producer's repository plus a check in CI.

Who owns the contract when the producer is a vendor feed or a legacy system nobody wants to touch?

The team that ingests it owns it. You cannot force a vendor into your CI, so the contract becomes a runtime assertion at the ingestion boundary: reject or quarantine records outside the declared domain, and version the contract when the vendor changes their feed. This is also the case where trace-side detection carries most of the load, because the producer will not announce the change.

How do contracts apply to unstructured retrieval corpora?

The structured metadata around the documents is where the contract lives: source system, document type, effective date, jurisdiction, access classification, and any field a retrieval filter uses. Chunking strategy and embedding model version belong in the contract too, because changing either alters what retrieval returns without changing a single document. Declare them as versioned terms and re-run the evaluation suite when they increment.

Does this slow producers down?

It adds a gate on Tier 1 and Tier 2 fields and a changelog entry on everything else. The tiering is what keeps it proportionate. If every field required consumer acknowledgement, producers would route around the process within a month.

Where do data contracts sit in an enterprise analytics platform?

They sit at the publication boundary of each governed domain, between the producer that owns a dataset and every downstream consumer. AWS data modernization guidance describes publishing one documented dataset per domain with named ownership, freshness expectations, and access controls [7], and the contract is how that description becomes enforceable rather than aspirational. In practice this means an analytics platform carries two artifacts per domain: the governed table itself, anchored to the Apache Iceberg table specification so snapshots and schema evolution are defined by the spec [1], and the contract file that declares the semantic and behavioral terms a registry cannot express. Reviewing the serving path against the AWS Well-Architected Generative AI Lens while the contract is being written keeps retrieval and inference concerns in the same design pass [8]. We build this boundary first in enterprise data and analytics platform engagements, because retrieval and agent layers added on top of an ungoverned domain inherit every ambiguity underneath them.

Where to Start

Go back to the rename that opened this article. Nothing in that incident required better monitoring. The null-rate check was working correctly; it just had no opinion about whether pending_review was a legitimate member of a set that four lines of routing code assumed was closed.

In your next working session, pick one Tier 1 field, write its semantics and enum domain into a contract file in the producer's repository, and add the check that fails a merge when the domain expands. Then inject a violation and watch the build go red. That single red build is worth more than a quarter of platform planning, because it is the first moment the contract exists in a form the organization cannot ignore.

Track the share of Tier 1 AI-consumed fields with a named steward and a declared contract version. If that number is not moving, nothing else in your data governance program is protecting the agent.

References

[1]Apache Software Foundation, "Apache Iceberg Table Specification," 2026. https://iceberg.apache.org/spec/

[2]OpenTelemetry, "Semantic Conventions for Generative AI Systems," 2026. https://opentelemetry.io/docs/specs/semconv/gen-ai/

[3]NIST, "AI 600-1: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile," 2024. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf

[4]NIST, "AI Risk Management Framework," 2026. https://www.nist.gov/itl/ai-risk-management-framework

[5]Yao et al., "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains," 2024. https://arxiv.org/abs/2406.12045

[6]Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" 2023 (still the reference pattern for executable outcome verification). https://arxiv.org/abs/2310.06770

[7]AWS, "Prescriptive Guidance: Database Migration Strategy," 2026. https://docs.aws.amazon.com/prescriptive-guidance/latest/strategy-database-migration/welcome.html

[8]AWS, "Well-Architected Generative AI Lens," 2026. https://docs.aws.amazon.com/wellarchitected/latest/generative-ai-lens/generative-ai-lens.html

[9]European Union, "Regulation (EU) 2024/1689 (Artificial Intelligence Act), Official Journal," 2024. https://eur-lex.europa.eu/eli/reg/2024/1689/oj

[10]Model Context Protocol, "Specification, 2025-06-18 revision," 2025. https://modelcontextprotocol.io/specification/2025-06-18

[11]AWS, "AWS Database Migration Service User Guide," 2026. https://docs.aws.amazon.com/dms/latest/userguide/Welcome.html

Article Summary

  1. 1Treat every field an AI system reads as a versioned contract term with a named producer owner
  2. 2Null-rate and row-count checks miss semantic breaks like a new enum value or reversed sign convention
  3. 3Pin analytical reads to the Apache Iceberg table spec so schema evolution is defined by the spec, not pipeline habit
  4. 4Gate producer merges in CI on contract compatibility, and replay archived traces when a contract version changes
  5. 5Score contract violations by consumer blast radius, not by table popularity

Ready to discuss this for your organization?

Talk to our team about implementing these approaches in your environment.

Get in Touch
Tactical Edge

AI workflows connected to the data, tools, and systems your teams use.

Washington, DC ยท United States

AWS PartnerAWS Advanced Tier Services Partner

AWS Generative AI Competency Partner

AWS Migration and Modernization Competency

Migration Services

Solutions

  • Agentic AI Systems
  • Agent Protocols (MCP/A2A)
  • AgentOps
  • Agent Governance
  • Moonshot Migrations
  • Cloud & Data
  • Amazon Quick
  • Amazon Connect
  • Document Automation
  • Industry Solutions
  • ISV Freedom Program

Platforms

  • Prospectory โ†—
  • Projectory โ†—
  • Monitory โ†—
  • Connectory โ†—
  • Greenway โ†—
  • Detectory โ†—

Services

  • Advisory & Strategy
  • Design & Engineering
  • Implementation
  • PoC & Pilot Programs
  • Agent Programs
  • Managed AI Operations
  • Governance & Compliance
  • AI Consulting

Company

  • About Us
  • Our Approach
  • AWS Partnership
  • Security
  • Demo Library
  • Events
  • Workshops
  • Insights & Resources
  • Careers
  • Contact

ยฉ 2026 Tactical Edge. All rights reserved.

Privacy PolicyTerms of ServiceAI PolicyCookie Policy