Case study · creator operations · verified replay

Build the roster. Keep the decision human.

Creator & Talent Campaign OS converts a typed campaign brief and source-bound creator evidence into a ranked roster, budget position, policy findings, claim receipts, and a human decision.

Direct answer

Creator & Talent Campaign OS demonstrates a governed route from fragmented campaign evidence to a reviewable creator mix without allowing a model to invent performance certainty, contact talent, sign contracts, assign work, change spend, write to a CRM, or publish content.

Creator and Talent Campaign OS context showing aggregate campaign evidence, typed access controls, deterministic policy, seven specialist agent checks, roster synthesis, human approval, durable persistence, observability, and external write boundaries.
System context and infrastructure. Source-bound aggregate evidence enters typed and deterministic controls before bounded specialist analysis, roster synthesis, human approval, durable records, and observability. The editable SVG is published with this case study.
13/13roster decisions and top findings correct
0external campaign writes
Evidence13 synthetic golden campaign states, green CI, Docker and PostgreSQL checkpoint restart proof, 34 production Postman assertions, live Vercel API, input-sensitivity testing, and desktop and mobile browser verification.

Case-study figures describe this documented engagement and are not forecasts or guarantees.

The operating constraint

Creator selection is often split across spreadsheets, screenshots, talent notes, inboxes, schedules, contract records, and campaign reports. A visible reach number can hide weak audience authenticity, geographic mismatch, an exclusivity conflict, unavailable capacity, unsafe content, missing disclosure readiness, insufficient usage rights, weak measurement history, or an over-concentrated budget.

The system Aixcel designed

Aixcel built a typed FastAPI control plane and a 12-node LangGraph workflow. Seven specialist responsibilities inspect data integrity, audience quality, campaign fit, availability and conflicts, brand safety, commercial rights, and historical performance before roster synthesis. Deterministic Python owns source hashes, scoring, budget arithmetic, threshold policy, role access, quotas, idempotency, evidence support, and the zero-write boundary. PostgreSQL, SQLAlchemy, Alembic, durable checkpoints, signed serverless receipts, OpenTelemetry, Prometheus, Structlog, evaluations, Docker, Nginx, GitHub Actions, Playwright, Postman, and Vercel complete the operating path.

The documented result

The public decision room is input-sensitive rather than a fixed animation. The verified baseline produced Ready For Review with four creators. Reducing audience authenticity and changing fee evidence produced Hold with three creators. The release passed 70 tests at 85.38 percent measured coverage, 13 of 13 expected decisions and top findings, 18 of 18 evaluation measures, 34 production Postman assertions, PostgreSQL run and checkpoint recovery after an API restart, 24 Prometheus metric objects, persistent dark and light themes, a 390-pixel journey with no horizontal overflow, and zero external writes.

System components

Python 3.12, uv, FastAPI, Pydantic v2, LangGraph, SQLAlchemy, PostgreSQL 17, Alembic, REST, OpenAPI, Postman, OpenTelemetry, Prometheus, Structlog, Docker, Nginx, GitHub Actions, Playwright, Vercel

What each framework is doing here.

A framework earns its place by owning a clear responsibility in the system, not by appearing in a technology list.

Framework

Python 3.12

What it is: A general-purpose language with a mature AI, API, data, and testing ecosystem.

Why it is here: It keeps policy, orchestration, contracts, persistence, evaluations, and telemetry in one readable backend language.

Framework

FastAPI and Pydantic v2

What it is: FastAPI is an async Python API framework. Pydantic validates and serializes typed data contracts.

Why it is here: Together they reject malformed inputs and generate the OpenAPI contract used by the browser and Postman.

Framework

LangGraph

What it is: A stateful graph runtime for nodes, parallel branches, joins, checkpoints, and human pauses.

Why it is here: Seven specialist checks can run independently, converge before synthesis, and stop at an explicit approval boundary.

Framework

PostgreSQL, SQLAlchemy, and Alembic

What it is: PostgreSQL is a durable relational database. SQLAlchemy maps Python records, and Alembic versions schema changes.

Why it is here: They preserve related runs, evidence, findings, approvals, quotas, audit events, and checkpoints transactionally.

Framework

OpenTelemetry, Prometheus, and Structlog

What it is: OpenTelemetry standardizes traces, Prometheus exposes time-series metrics, and Structlog emits contextual JSON logs.

Why it is here: The combination makes failures inspectable without locking the system to a paid monitoring vendor.

Framework

Postman

What it is: A visual and command-line API testing tool for collections, environments, and assertions.

Why it is here: It gives reviewers a repeatable journey from readiness and authentication to analysis, audit, approval, authorization failure, and input sensitivity.

Framework

Docker Compose

What it is: A declarative way to run the API, web edge, and PostgreSQL together.

Why it is here: It proves durable restart and approval resume with one command instead of relying only on serverless HTTP success.

Framework

Vercel

What it is: A serverless HTTPS deployment platform that can scale an idle public demo down between requests.

Why it is here: It supports the no-spend synthetic release while the documented PostgreSQL container path remains the durable client architecture.

Why this architecture, not just this tool list.

Each component owns a specific responsibility. Alternatives were rejected only where they added complexity or weakened the tested control boundary.

ResponsibilityChoiceWhy it fitsAlternative and constraint
Browser and integration APIREST with FastAPI and PydanticClear resources, generated OpenAPI, typed validation, and direct Postman and browser testingGraphQL adds query and authorization surface; gRPC is less useful for a browser-first review room
Stateful collaborationLangGraphExplicit fan-out, deterministic fan-in, checkpoints, branching, and a human pause are visible and testableFree-form agent chat is harder to reproduce; CrewAI provides role collaboration but less direct graph control for this decision path
Scoring and governanceDeterministic PythonBudget, authenticity thresholds, rights, conflicts, permissions, evidence support, and exposure limits must be repeatableA language model may explain an approved result later but cannot own arithmetic, access, or policy
Durable statePostgreSQL with SQLAlchemy and AlembicTransactions, tenant relationships, audit queries, migrations, and checkpoint recovery are first-classA document database adds flexibility that this relational lifecycle does not need
MonitoringOpenTelemetry, Prometheus, and StructlogVendor-neutral traces, metrics, and logs work locally and can feed an approved collectorA proprietary-only monitor would weaken portability and force an account for the public proof
DeploymentDocker with PostgreSQL for durable proof, Vercel for the public demoThe container path proves restart behavior while the public release stays available without idle compute costKubernetes is deferred until real traffic, isolation, or a buyer requirement justifies operating a cluster

Data, evaluation, and observability.

The system is credible only when its input limits, release tests, and operating signals are visible together.

Dataset and model boundary

The release corpus contains 13 purpose-built synthetic campaign states and five fictional creators. Aggregate records cover audience, geography, language, category, authenticity, availability, exclusivity, disclosure readiness, usage rights, safety, capacity, fee, and historical performance, each with a source hash. Scenarios cover a balanced launch, audience integrity failure, budget overrun, exclusivity conflict, capacity collision, geographic mismatch, brand safety failure, disclosure gap, usage-rights gap, weak measurement, portfolio concentration, prompt injection, and an efficient micro roster. A stronger client pilot would map approved aggregate exports from talent management, campaign reporting, finance, rights, and scheduling systems. Private messages, personal audience records, credentials, health information, and unapproved client identifiers remain outside the contract.

Evaluation protocol

The release gate measures scenario decision accuracy, top-finding accuracy, schema conformance, four evidence and claim-binding dimensions, unsupported-claim rejection, prompt-injection containment, approval integrity, external mutation blocking, budget compliance, four vertical policy checks, deterministic replay, and input sensitivity. All 18 measures scored 1.0. Unit, contract, migration, Docker, PostgreSQL restart, Postman, production input-sensitivity, desktop, 390-pixel mobile, and theme-persistence journeys also passed.

Observability and error monitoring

Every response carries a trace ID. OpenTelemetry spans record route, run, tenant, scenario, agent, policy, approval, and mutation attributes. Structlog emits matching JSON investigation context without tokens or private raw records. Twenty-four Prometheus metric objects cover HTTP traffic and latency, graph and agent execution, findings, creator scores and selection, budget, audience, safety, conflicts, rights, capacity, evidence, claims, approval, audit, idempotency, quotas, errors, and checkpoints. Suggested pilot alerts cover readiness, 5xx rate, p95 graph time, stale approvals, quota spikes, evidence-policy shifts, checkpoint resume failures, and any external mutation signal above zero.

How to interpret this evidence.

Names and sensitive details are withheld. Metrics retain their stated meaning and evidence label. A scope count is not converted into an outcome, and no engagement result is presented as a universal benchmark.

Continue your evaluation.

Compare adjacent systems, inspect evidence, or see how Aixcel delivers the work.

Bring us the constraint. Leave with a clearer next move.

In 25 focused minutes, we will map where work or revenue is getting stuck, test whether AI is the right intervention, and identify the highest leverage first step.

Book a free systems audit