Auditability and evaluation
of agentic workflows.

Overview
When AI workflows broke in production, builders couldn't see what went wrong, and quality checks before shipping were too slow and manual to stick. I redesigned analytics and continuous evaluation into one org-wide system for watching agent health.
Company
StackAI
Feature
Workflow Analytics & Evaluator
Year
2026
Impact
Usage grew 5.5× with deeper user engagement. Runs drill downs drove 22,000+ views in six months.

PROCESS

Personas and user journeys for analytics and evaluation Benchmark analytics competitive research Evaluator user flow from sandbox to production Design ideation with PM and engineering

1. Personas and problem space

Executive sponsors need ROI proof, admins need production health, builders need to debug and iterate, and StackAI needs platform usage insight. Analytics was split across two surfaces with no org-wide view, buried run logs, and timezone bugs. The evaluator was barely used: manual test uploads and deployment friction, while competitors had moved to continuous, trace-native evaluation.
Personas and user journeys for analytics and evaluation

2. Benchmark

Competitive research across Laminar, Langfuse, and Arize Phoenix. Mapped the opportunity space in FigJam, ran design iterations across global and project-level views, and validated with stakeholders before engineering handoff.
Benchmark analytics competitive research

3. User journeys and eval flow

Mapped how builders run workflows, debug failures, and loop back through analytics. Designed the end-to-end eval journey from sandbox testing through production: build, publish, monitor, and continuous evaluation with human-in-the-loop gates.
Evaluator user flow from sandbox to production

4. Ideation with PM + Eng

Collaborative design sessions exploring UI directions across node editors, data tables, and evaluation workflows. Sketched run-level detail views, trace timelines, and error clustering patterns before converging on a shared concept.
Design ideation with PM and engineering

SOLUTION

Improved analytics. Production failures were hard to debug because analytics lived across split surfaces with buried run logs and no org-wide view. We redesigned analytics around caller context and feedback filtering, with debugging as the primary job, so builders and admins can see what broke and who was affected without hopping tools.

Analytics overview with run metrics, trend chart, and project table
Adoption analytics with runs by group, users, and usage leaderboard
Model usage table grouped by provider and model
Run logs table with status, caller, latency, and output

Analytics view a la Temporal. Stakeholders needed success rates, latency, and error trends in one place, not scattered project dashboards. Inspired by Temporal-style workflow health, the org-wide view makes production status scannable over time so platform and admin personas can spot risk before it becomes a support fire drill.

Org-wide analytics view inspired by Temporal workflow metrics

Agent evaluator. Manual test uploads and deployment friction meant evaluation barely happened, while competitors moved to continuous, trace-native checks. We connected analytics, evaluator, and test datasets through one shared data model so a production run can move into evaluation and back into iteration without leaving the product loop.

Evaluator table with LLM judges, criteria, sample rate, and production toggles
Edit evaluator criteria, scoring thresholds, and audience settings
Evaluator playground for testing judges against sample runs
Evaluator run results with scores and suggested changes

INSIGHTS

Two sides of the same coin

Analytics and evaluation are interconnected. A run can be viewed in analytics, sent to an evaluator, and used to build a test dataset, the two systems share a data model by design.

Adoption ≠ debugging

Executive sponsors and builders need separate views of the same data, not just more filters.

Friction kills good ideas

The evaluator existed but went unused because deploying it required infrastructure work. The new model makes signals zero-configuration.

ADLC makes evaluation strategic

When evals are deployment gates, defining "good" becomes a required step before shipping, not a retrospective.

OTHER PROJECTS