Guardrails for AI-Generated Python Code: A pytest + Hypothesis Setup That Catches What Review Misses
AI assistants are very good at writing Python that looks right. It's typed, it's formatted, it comes with tests. The problem is that the tests are usually written by the same model, from the same assumptions, so they pass for the same reasons the code might be wrong.
I've spent a lot of time with Python code that processes sensor and telemetry data, and at some point I stopped trusting review alone to catch this. Here's the setup I use now. It's three layers, and each one catches a different kind of mistake.
The Failure Mode
A typical example: you ask an assistant for a function that converts power readings to watts. You get this:
# telemetry.py
from decimal import Decimal
UNIT_FACTORS = {"W": Decimal("1"), "kW": Decimal("1000"), "MW": Decimal("1000000")}
def to_watts(value: Decimal, unit: str) -> Decimal:
if value < 0:
raise ValueError("power reading cannot be negative")
try:
return value * UNIT_FACTORS[unit]
except KeyError:
raise ValueError(f"unknown unit: {unit}") from None
This version is fine. The versions that aren't fine look almost identical: a float instead of Decimal, a missing negative check, "kw" instead of "kW". Example-based tests with three hand-picked values rarely catch these. Properties do.
Layer 1: Property-Based Tests with Hypothesis
Instead of asserting specific outputs, assert things that must always be true:
# test_telemetry.py
from decimal import Decimal
import pytest
from hypothesis import given, strategies as st
from telemetry import UNIT_FACTORS, to_watts
readings = st.decimals(
min_value=0,
max_value=10**6,
places=3,
allow_nan=False,
allow_infinity=False,
)
@given(value=readings, unit=st.sampled_from(list(UNIT_FACTORS)))
def test_larger_units_never_produce_smaller_values(value, unit):
assert to_watts(value, unit) >= to_watts(value, "W")
@given(value=readings)
def test_kw_round_trip_is_exact(value):
assert to_watts(value, "kW") / 1000 == value
@given(
value=st.decimals(
max_value=Decimal("-0.001"),
allow_nan=False,
allow_infinity=False,
)
)
def test_negative_readings_are_rejected(value):
with pytest.raises(ValueError):
to_watts(value, "W")
The round-trip test fails immediately if someone "simplifies" the function to use floats. The negative-reading test fails if the guard disappears in a refactor. Neither needs you to guess the breaking input.
I write properties by hand, not with the assistant. That's the point: the tests encode what a human believes must be true, independently of what the model generated.
Layer 2: Schema Snapshots for Data Contracts
AI refactors love to "improve" models: an int becomes float, an optional field becomes required. On a system talking to thousands of devices, that's a silent breaking change. I snapshot the JSON schema of every Pydantic model that crosses a boundary:
# test_contracts.py
import json
from pathlib import Path
from models import Reading # your Pydantic v2 model
SNAPSHOT = Path(__file__).parent / "snapshots" / "reading.schema.json"
def test_reading_schema_is_stable():
current = Reading.model_json_schema()
expected = json.loads(SNAPSHOT.read_text())
assert current == expected, "Reading schema changed: update the snapshot deliberately"
When the schema changes on purpose, the developer updates the snapshot in the same PR, and the diff makes the contract change visible to reviewers. When it changes by accident, CI stops it.
Layer 3: Mutation Testing, but Only Where It Matters
Mutation testing (I use mutmut) changes your code in small ways and checks whether any test fails. If a mutant survives, your tests don't really cover that logic, no matter what the coverage report says. It's slow, so I don't run it on every PR. I limit it to critical modules and run it nightly:
# pyproject.toml
[tool.mutmut]
paths_to_mutate = ["src/telemetry.py"]
mutmut run
mutmut results
(Config options differ between mutmut versions, so check the docs for yours.)
Surviving mutants go into the backlog as test gaps.
Wiring It into CI
Hypothesis profiles keep local runs fast and CI runs thorough:
# conftest.py
from hypothesis import settings
settings.register_profile("dev", max_examples=50)
settings.register_profile("ci", max_examples=300)
# .github/workflows/tests.yml
name: tests
on: [pull_request]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: pip install -r requirements-dev.txt
- run: pytest --hypothesis-profile=ci
Who Writes What
The layers only work if the split between human and assistant is explicit. Mine looks like this:
- I write the properties and schema snapshots by hand.
- The assistant writes the implementation and example-based tests.
I also keep a short rules file in the repo that the assistant reads (a CLAUDE.md or equivalent):
Decimal for all power and money values
no sync DB calls inside async def
never edit files under snapshots/
It doesn't replace the tests, but it reduces how often they fail.
Checklist for Your Repo
- Pick two or three modules where a wrong value costs real money or safety.
- Write three to five properties for each, by hand.
- Snapshot the schemas of every model that crosses a service boundary.
- Run mutation testing nightly on those modules only.
- Add a rules file for your assistant based on the bugs the layers catch.
What This Doesn't Solve
- It won't catch wrong architecture. A well-tested function in the wrong place is still in the wrong place. That still needs a senior reviewer.
- Good properties are hard to write. Expect the first few to take longer than the code they test.
- Hypothesis adds CI time. In my experience, 300 examples per property adds seconds, not minutes, but heavy fixtures can change that.
For me the trade-off has been worth it on systems where a wrong number in production means a wrong invoice or a wrong charging decision. On a throwaway prototype, it probably isn't.
What does your setup look like? I'm curious whether others gate AI-generated code differently, especially around async code.
If you'd rather not build this from scratch: Boldare, a Python software house in Poland, runs exactly this kind of guardrail setup on production backends for IoT and energy systems, with senior review on every merge. Worth a look if you need a team that takes AI-generated code seriously.
Comments
No comments yet. Start the discussion.