Case study · creative intelligence · verified replay
Know what worked. Know why it might have.
Creative Learning OS checks whether a campaign comparison can support a lesson before it drafts the next test. It separates real signal from attribution, audience, placement, sample, lag, funnel-quality, and fatigue confounds.
Creative Learning OS demonstrates a governed route from aggregate campaign evidence to a reviewable creative learning without allowing a model to invent causal certainty, hide limitations, publish content, or change media spend.

Case-study figures describe this documented engagement and are not forecasts or guarantees.
The operating constraint
Creative teams can call an asset a winner because CTR rose even when attribution windows, audience warmth, placement, sample size, conversion maturity, or creative variables changed at the same time. A false lesson then enters the next brief and causes the team to spend more money reproducing noise.
The system Aixcel designed
Aixcel built a typed FastAPI control plane and an 11-node LangGraph workflow. Six analysis responsibilities inspect data integrity, creative taxonomy, measurement validity, normalized funnel performance, fatigue, and the public baseline before synthesis. Deterministic Python owns source hashing, metric formulas, sample and comparison gates, evidence policy, exposure limits, access, quotas, and the zero-mutation boundary. PostgreSQL, SQLAlchemy, Alembic, durable checkpoints, OpenTelemetry, Prometheus, structured logs, evaluations, Docker, Nginx, GitHub Actions, Playwright, Postman, and Vercel complete the operating path.
The documented result
The public workspace is input-sensitive rather than a fixed animation. The verified browser journey produced Scale with positive conversion evidence, then changed to Hold and reported lift as unavailable when the conversion comparator was set to zero. Mixed attribution fails closed. The release passed 68 tests, 82.76 percent measured coverage, 13 of 13 expected top findings and decisions, 18 of 18 evaluation measures at target, 34 production Postman assertions, PostgreSQL checkpoint recovery across an API restart, persistent dark and light themes, 20 Prometheus metric objects, and zero external mutations.
System components
Python 3.12, FastAPI, Pydantic v2, LangGraph, SQLAlchemy, PostgreSQL 17, Alembic, REST, OpenAPI, Postman, OpenTelemetry, Prometheus, Docker, Nginx, GitHub Actions, Playwright, Vercel
Why this architecture, not just this tool list.
Each component owns a specific responsibility. Alternatives were rejected only where they added complexity or weakened the tested control boundary.
| Responsibility | Choice | Why it fits | Alternative and constraint |
|---|---|---|---|
| Browser and integration API | REST with FastAPI and Pydantic | Clear resources, generated OpenAPI, typed validation, simple Postman and browser testing | GraphQL adds query complexity; gRPC does not fit a browser-first review surface |
| Stateful collaboration | LangGraph | Explicit fan-out, deterministic fan-in, checkpoints, branching, and a human pause are visible and testable | Free-form agent chat and CrewAI are less direct for this control-heavy workflow |
| Measurement policy | Deterministic Python | Attribution, sample, formulas, permissions, and exposure limits must be repeatable | An LLM may explain an approved result later but cannot own arithmetic or policy |
| Durable state | PostgreSQL with SQLAlchemy and Alembic | Transactions, tenant relationships, audit queries, migrations, and checkpoint recovery are first-class | MongoDB offers flexibility that this relational lifecycle does not need |
| Monitoring | OpenTelemetry, Prometheus, and structured logs | Vendor-neutral signals work locally and can connect to Langfuse, Grafana, or another approved collector | A proprietary-only monitor would weaken portability and the zero-cost replay |
| Deployment | Docker plus Nginx for portable release, Vercel for the public demo | The container path proves a private API, database, health checks, and durable state while the public site scales to zero | Kubernetes is deferred until a real target needs cluster scheduling and horizontal scale |
Data, evaluation, and observability.
The system is credible only when its input limits, release tests, and operating signals are visible together.
Dataset and model boundary
The regression corpus contains 13 synthetic aggregate portfolios covering valid lift, mixed attribution, fatigue, low sample, click-to-conversion conflict, placement and audience confounds, missing taxonomy, conversion lag, hostile caption instructions, multi-variable tests, stable control, and source receipt tampering. A separate CC BY 4.0 UCI Facebook Metrics artifact contains 500 historical rows, with 400 sequential training rows and 100 holdout rows. It has no geography, Florida cohort, influencer identity, or sales-outcome contract, so the public ridge model is disclosed as a low-confidence historical range, not a promised forecast.
Evaluation protocol
The release gate measures condition detection, top-finding accuracy, decision accuracy, schema conformance, five evidence-binding dimensions, claim taxonomy, prompt-injection resistance, measurement policy, mutation safety, exposure bounds, source hashes, approval integrity, zero-cost replay, and model-scope disclosure. All 18 measures scored 1.0 across 13 scenarios. Unit, contract, migration, Docker, PostgreSQL restart, Postman, input-sensitivity, desktop, 390-pixel mobile, and theme-persistence journeys must also pass.
Observability and error monitoring
Every response carries a trace ID. Structured JSON logs, OpenTelemetry export, a Prometheus endpoint, lifecycle audit receipts, and 20 named metric objects cover traffic, latency, run state, node execution, finding categories, human decisions, observed lift, budget exposure, policy outcomes, idempotency, quotas, safe errors, checkpoint resumes, sample blocks, fatigue, attribution mismatch, and public baseline use. Initial pilot alerts are documented for readiness, 5xx rate, p95 graph time, hash mismatch, approval resume failure, policy shift, fatigue shift, and model drift.
How to interpret this evidence.
Names and sensitive details are withheld. Metrics retain their stated meaning and evidence label. A scope count is not converted into an outcome, and no engagement result is presented as a universal benchmark.
Continue your evaluation.
Compare adjacent systems, inspect evidence, or see how Aixcel delivers the work.
Bring us the constraint. Leave with a clearer next move.
In 25 focused minutes, we will map where work or revenue is getting stuck, test whether AI is the right intervention, and identify the highest leverage first step.
Book a free systems audit