Case study · creative intelligence · verified replay

Know what worked. Know why it might have.

Creative Learning OS checks whether a campaign comparison can support a lesson before it drafts the next test. It separates real signal from attribution, audience, placement, sample, lag, funnel-quality, and fatigue confounds.

Direct answer

Creative Learning OS demonstrates a governed route from aggregate campaign evidence to a reviewable creative learning without allowing a model to invent causal certainty, hide limitations, publish content, or change media spend.

Creative Learning OS system context showing aggregate campaign evidence, typed contracts, deterministic controls, LangGraph collaboration, persistence, human approval, observability, and external mutation boundaries.
System context and infrastructure. Aggregate evidence enters typed and deterministic controls before bounded analysis, evidence policy, human approval, durable records, and observability. The editable SVG is published with this case study.
13/13top findings and decisions correct
0automatic platform mutations
Evidence13 synthetic golden scenarios, a licensed public historical baseline, green CI, Docker and PostgreSQL checkpoint proof, 34 production Postman assertions, live Vercel API, input-sensitivity testing, and desktop and mobile browser verification.

Case-study figures describe this documented engagement and are not forecasts or guarantees.

The operating constraint

Creative teams can call an asset a winner because CTR rose even when attribution windows, audience warmth, placement, sample size, conversion maturity, or creative variables changed at the same time. A false lesson then enters the next brief and causes the team to spend more money reproducing noise.

The system Aixcel designed

Aixcel built a typed FastAPI control plane and an 11-node LangGraph workflow. Six analysis responsibilities inspect data integrity, creative taxonomy, measurement validity, normalized funnel performance, fatigue, and the public baseline before synthesis. Deterministic Python owns source hashing, metric formulas, sample and comparison gates, evidence policy, exposure limits, access, quotas, and the zero-mutation boundary. PostgreSQL, SQLAlchemy, Alembic, durable checkpoints, OpenTelemetry, Prometheus, structured logs, evaluations, Docker, Nginx, GitHub Actions, Playwright, Postman, and Vercel complete the operating path.

The documented result

The public workspace is input-sensitive rather than a fixed animation. The verified browser journey produced Scale with positive conversion evidence, then changed to Hold and reported lift as unavailable when the conversion comparator was set to zero. Mixed attribution fails closed. The release passed 68 tests, 82.76 percent measured coverage, 13 of 13 expected top findings and decisions, 18 of 18 evaluation measures at target, 34 production Postman assertions, PostgreSQL checkpoint recovery across an API restart, persistent dark and light themes, 20 Prometheus metric objects, and zero external mutations.

System components

Python 3.12, FastAPI, Pydantic v2, LangGraph, SQLAlchemy, PostgreSQL 17, Alembic, REST, OpenAPI, Postman, OpenTelemetry, Prometheus, Docker, Nginx, GitHub Actions, Playwright, Vercel

Why this architecture, not just this tool list.

Each component owns a specific responsibility. Alternatives were rejected only where they added complexity or weakened the tested control boundary.

ResponsibilityChoiceWhy it fitsAlternative and constraint
Browser and integration APIREST with FastAPI and PydanticClear resources, generated OpenAPI, typed validation, simple Postman and browser testingGraphQL adds query complexity; gRPC does not fit a browser-first review surface
Stateful collaborationLangGraphExplicit fan-out, deterministic fan-in, checkpoints, branching, and a human pause are visible and testableFree-form agent chat and CrewAI are less direct for this control-heavy workflow
Measurement policyDeterministic PythonAttribution, sample, formulas, permissions, and exposure limits must be repeatableAn LLM may explain an approved result later but cannot own arithmetic or policy
Durable statePostgreSQL with SQLAlchemy and AlembicTransactions, tenant relationships, audit queries, migrations, and checkpoint recovery are first-classMongoDB offers flexibility that this relational lifecycle does not need
MonitoringOpenTelemetry, Prometheus, and structured logsVendor-neutral signals work locally and can connect to Langfuse, Grafana, or another approved collectorA proprietary-only monitor would weaken portability and the zero-cost replay
DeploymentDocker plus Nginx for portable release, Vercel for the public demoThe container path proves a private API, database, health checks, and durable state while the public site scales to zeroKubernetes is deferred until a real target needs cluster scheduling and horizontal scale

Data, evaluation, and observability.

The system is credible only when its input limits, release tests, and operating signals are visible together.

Dataset and model boundary

The regression corpus contains 13 synthetic aggregate portfolios covering valid lift, mixed attribution, fatigue, low sample, click-to-conversion conflict, placement and audience confounds, missing taxonomy, conversion lag, hostile caption instructions, multi-variable tests, stable control, and source receipt tampering. A separate CC BY 4.0 UCI Facebook Metrics artifact contains 500 historical rows, with 400 sequential training rows and 100 holdout rows. It has no geography, Florida cohort, influencer identity, or sales-outcome contract, so the public ridge model is disclosed as a low-confidence historical range, not a promised forecast.

Evaluation protocol

The release gate measures condition detection, top-finding accuracy, decision accuracy, schema conformance, five evidence-binding dimensions, claim taxonomy, prompt-injection resistance, measurement policy, mutation safety, exposure bounds, source hashes, approval integrity, zero-cost replay, and model-scope disclosure. All 18 measures scored 1.0 across 13 scenarios. Unit, contract, migration, Docker, PostgreSQL restart, Postman, input-sensitivity, desktop, 390-pixel mobile, and theme-persistence journeys must also pass.

Observability and error monitoring

Every response carries a trace ID. Structured JSON logs, OpenTelemetry export, a Prometheus endpoint, lifecycle audit receipts, and 20 named metric objects cover traffic, latency, run state, node execution, finding categories, human decisions, observed lift, budget exposure, policy outcomes, idempotency, quotas, safe errors, checkpoint resumes, sample blocks, fatigue, attribution mismatch, and public baseline use. Initial pilot alerts are documented for readiness, 5xx rate, p95 graph time, hash mismatch, approval resume failure, policy shift, fatigue shift, and model drift.

How to interpret this evidence.

Names and sensitive details are withheld. Metrics retain their stated meaning and evidence label. A scope count is not converted into an outcome, and no engagement result is presented as a universal benchmark.

Continue your evaluation.

Compare adjacent systems, inspect evidence, or see how Aixcel delivers the work.

Bring us the constraint. Leave with a clearer next move.

In 25 focused minutes, we will map where work or revenue is getting stuck, test whether AI is the right intervention, and identify the highest leverage first step.

Book a free systems audit