Book a scoping call

OARUX // agent UX research, end to end

Propelling the human-agent experience.

Ship an agent real users actually trust, across its whole lifecycle: discover real demand, design with real users, validate behavior against ground truth, then operate and benchmark every release. Eval System Build is the keystone, the calibrated judge that becomes your standing eval asset.

discover: signal validate: pass operate: drift regression: fail
The trust gap

The hard part isn't shipping an agent. It's knowing real users will trust it, release after release.

A trustworthy agent is a research problem at every stage: which demand is real, what users do with a working prototype, whether behavior holds up against ground truth, and whether it stays honest once it's live. Your team already runs golden sets, judges, and tracing, a competent and necessary stack. The one input it can't generate from inside its own walls is criteria grounded in how real users actually behave. That's the sharp end, and it's where we work.

01

The in-house judge is an echo chamber

An LLM-as-a-judge can only grade against the criteria written for it, and in-house those criteria inherit the team's own model of “good.” An agent can clear the internal judge while real users quietly churn. The loop optimizes a proxy that was never checked against real user behavior.

02

A score, never the why

Benchmarks and thumbs-up/thumbs-down feedback tell you a session scored low. They do not tell you why. The signals that are easy to capture (latency, token cost, semantic similarity) are not the friction axes that actually move retention.

03

Likert is noise to an engineer

Where human signal does get collected, it is usually a coarse rating. The difference between a 3 of 5 and a 4 of 5 trust score gives an engineer nothing to change, and a one-off study goes stale the moment the next model swap ships.

What we do

One lifecycle. Five ways in.

Five offerings along a single arc: Discover, Design, Validate, Operate, Benchmark. Land on a sharp, self-contained piece and grow it into a standing eval program. Every engagement hands back deployable infrastructure with stated reliability, not a deck.

  1. Discover

    Signal

    Map where real demand and friction live in your domain, using public and community data competitors aren't watching, so the roadmap is grounded before the first dollar is spent.

  2. Design

    Participatory Agent Design

    Real target users shape the agent on a working prototype, so its behavior is right while it is still cheap to change, instead of the gaps surfacing after launch.

  3. Operate

    Ongoing Evals

    Auto-score live traces against your calibrated judge and catch regressions before users feel them, as the model, prompts, and user base shift.

  4. Benchmark

    ROAB Benchmark

    The Reference-Outcome Agent Benchmark (ROAB) answers “is v2 actually better for real users?” and “are we ahead of the competitor?” on task completion, with real users and ground-truth reference outcomes, not vibes.

Most teams start with the keystone or a self-contained benchmark, then grow into the retained program. We will scope the right entry point with you.

Book a scoping call
Inside the keystone

Open, axial, rubric, deploy.

Eval System Build is the keystone of the lifecycle, and this is how it works: inductive grounded theory applied to AI evals. The sequence is load- bearing: reliability is proven before anything becomes a metric, and the metric is calibrated before anything deploys. Each gate is a hard stop.

  1. Step 1 Inductive discovery

    Open coding

    Unconstrained labs with real users on their real goals, no script. Two or more expert raters tag emergent friction line by line, grounded in observed behavior rather than a pre-built dimension list.

    GATETarget: Fleiss κ ≥ 0.80 inter-rater reliability before anything proceeds.

  2. Step 2 Derive the axes

    Axial coding

    Cluster the reliable codes into the category-specific behavioral axes that recur for this task type, then test which axes associate with abandonment. Discovered from the data, not imported from a fixed 8 to 12 dimension template.

    GATEAxes named behaviorally, statistically supported, reported with their limits.

  3. Step 3 Binary, never Likert

    Rubric design

    Each axis becomes a binary or structured check an engineer can act on. Calibrate the LLM-judge against a human-coded golden set via a confusion matrix until it agrees with expert labels.

    GATEGate: judge TPR ≥ 0.80 and TNR ≥ 0.80 per criterion against the held-out golden set; 0.90/0.90 is the default target, not the gate.

  4. Step 4 The judge that runs forever

    Deploy

    Install the calibrated judge into your Langfuse (primary, self-hostable in your VPC) or Arize Phoenix, auto-score live traces, and alert on drops. Then we hand back the golden set and the runbook, and leave.

    GATELive scores reproduce golden-set agreement; every commit and model swap is re-checked.

These targets are the bars every engagement clears before a number ships. The numbers you receive are reported as measured, never rounded up to the target.

Human-in-the-loop

Real users produce the ground truth.

Every axis and every rubric traces back to behavior observed from real participants in moderated labs, coded line by line by two or more expert raters. Those people are the instrument your judge is calibrated against.

Research participant in a moderated usability lab
P-01
Research participant in a moderated usability lab
P-02
Research participant in a moderated usability lab
P-03
Research participant in a moderated usability lab
P-04
Research participant in a moderated usability lab
P-05
Research participant in a moderated usability lab
P-06

Real humans in the loop // ground-truth reference outcomes // multi-turn interaction quality

Illustrative example

No client engagement is represented here. OARUX has not yet shipped a delivered client engagement as public proof — every number below is self-generated on a constructed example, to show what the method looks like executed correctly, not to claim it has been run for a client.

Worked example // open coding to a calibrated judge

The method, executed correctly, on a worked example.

A support agent handling subscription cancellation and downgrade requests, run through the pipeline end to end. Reliability was checked two ways before any judge-calibration budget was spent — not just whether raters agreed overall, but whether they agreed enough on each specific class to make the judge's target reachable at all. Each rubric criterion was then tuned on one split of the golden set and measured, once, on a separate held-out split — reported per criterion below, never pooled into one headline pair, because a pooled average is exactly what hides a criterion that still needs work.

Stage 1 — reliability
0.84

Fleiss κ Inter-rater reliability across three raters, open-coding the same transcripts independently. Cleared the 0.80 gate before any code became a rubric item.

Clearing κ isn't the whole check. Before any judge-calibration budget is spent, an internal check also asks whether the raters agreed enough on each specific class — ppos and pneg, positive- and negative-specific rater agreement — to make the judge's target reachable at all: a judge can never be more consistent about a class than the humans who labelled it were. This runs on a two-rater double-labelled subset (Cohen's κ here, not Fleiss, since ppos/pneg are pairwise by definition, distinct from the three-rater Fleiss κ above). It is a precondition on spending, not a published client metric — ppos and pneg are never measured or reported to the same standard as the judge's rates below, and no rate from them appears alongside those rows.

Candidate criterion: correctly tags the customer's true cancellation reason Screened out — never reaches Stage 2

One in five positive labels on this criterion were contested between the two raters, and Cohen's κ 0.60 fails the 0.80 Stage-1 gate outright. That disagreement caps what the judge could ever be asked for here — asking for a TNR of 0.90 on this criterion would ask the machine to be more consistent than the people who labelled it were, which isn't ambitious, it's incoherent. Screened out and rewritten before it is ever scored against a single judge — an internal gate result that stopped the work, not a metric that would sit beside the published rates below.

Stage 2 — judge calibration

Every criterion is tuned on one split of the golden set, then measured exactly once on a second, genuinely held-out split. Only the held-out numbers are ever reported as the gate result — a criterion never grades its own tuning data. Certification itself reads the wider cluster-bootstrap 95% interval, resampled over participants, not the tighter Wilson interval shown below — Wilson assumes independent judgements and is the optimistic approximation. The rates here are chosen with enough headroom that they would still certify under the wider interval too.

  • Grounds retention offer in the actual plan terms Certified
    TPR 0.95 (held-out, n=100) TNR 0.96 (held-out, n=100)
  • Does not fabricate a discount or policy exception Certified
    TPR 0.94 (held-out, n=100) TNR 0.93 (held-out, n=100)
  • Escalates when it lacks authority to resolve the request Failed — sent back
    Tuning split (n=50/50) TPR 0.94 TNR 0.88

    Looked like a clear on both rates.

    Held-out split (n=50/50) TPR 0.90 TNR 0.42

    TPR barely moves. TNR collapses.

    Held-out TNR's 95% CI is [0.29, 0.56] — the upper bound stays well below 0.80. Certifying needs the lower bound above 0.80; failing definitively needs only the upper bound below it, which this clears easily. Cycles back through rubric or judge-prompt iteration — more labelling would not fix it.

  • Confirms the cancellation before closing the chat Unmeasurable
    TPR 0.91 (held-out, n=100) TNR unmeasurable (12 negatives, held-out)

    Only 12 negative-class judgements turned up in the held-out set — below the 20-judgement floor. Reported as unmeasurable, not estimated.

Certified — the point estimate clears 0.80 and the interval's lower bound does too (cluster-bootstrap in delivered work): the measurement can be trusted. Failed — the point estimate itself misses 0.80; more labelling does not fix this, only rewriting the criterion or the judge prompt does. Unmeasurable — fewer than 20 judgements in a class; not a pass, not a fail, just not enough data to say either.

Why the headline is never pooled

Take five typical criteria at TNR 0.96 (48/50 negative-class judgements each) and one broken one at TNR 0.30 (15/50) — a criterion that only catches 30% of the failure it exists to detect. Pool all six into one negative-class denominator and the headline reads TNR 0.85 on 255/300 judgements, Wilson lower bound 0.8052 — comfortably "certified." The pooled number is arithmetically real, and it would ship a judge that misses 70% of one failure mode while calling itself calibrated. This is why every rate on this page is reported per criterion, never pooled — and why a rubric that has never failed a criterion in an engagement isn't evidence the rubric is good. It's evidence the gate isn't real.

Confusion matrix, criterion 1 only illustrative
Present Absent
Judge: present 95 4
Judge: absent 5 96

The arithmetic behind row 1 above ("grounds retention offer in the actual plan terms") only — not a pooled aggregate across the rubric. 95% Wilson CI: TPR [0.89, 0.98], TNR [0.90, 0.98] — both roughly ten points clear of the 0.80 gate on their lower bound, chosen with that much headroom deliberately, because Wilson assumes independence and is optimistic. Delivered work reports the wider cluster-bootstrap interval, resampled over participants; it is shown here as a Wilson approximation since there are no real participants to bootstrap. There is no client behind these counts.

Note: these establish judge-to-human agreement on a worked example, a calibration result, not a business-outcome claim. Where OARUX links rubric metrics to retention or satisfaction, those links are reported as association. Only a controlled A/B test on a deployed change can establish that improving a metric drives an outcome.

The founder

Twenty-five years of UX research, rebuilt for the agent era.

Trust has been the thread of Andy Hay's work for twenty-five years — from an MSc thesis at University College London that became a patented trusted-computing interface, designed at HP Labs' Trusted Platforms Group, to OARUX, where the same question of when to trust the machine is now the premise. Across those years he has built research teams, products, and the systems behind them, with a research practice spanning AI/ML, cloud, and data and analytics — the exact surfaces agent teams build on today. OARUX points that rigor at something your engineering team can deploy.

Classic UX research ends in a deck. OARUX ends in a calibrated judge running in your observability layer.

Earlier, as a UX Research Lead at Microsoft (2002–2008), he was named on a product patent for Windows Server Update Services and earned the company's Gold Star award for the work. He went on to co-found User Research International (URI), scaling it from two people to more than a hundred over seventeen years before exiting in 2025. URI's engagements included foundational research for Microsoft Azure and Google Cloud Platform, and reached seven of the world's ten largest technology companies — the same product, engineering, and data teams OARUX is built for today. Beyond client work he built a 76,000-participant research panel and took Panel Pro, a SaaS platform, from concept to production: research delivered as working systems. His MSc in Ergonomics & HCI is from UCL, his BSc in Psychology from Goldsmiths, University of London.

Data handling Your data stays yours. The calibrated judge and its scoring run inside your own observability layer, self-hostable in your VPC on Langfuse or Arize Phoenix. Engagements are NDA-friendly, and this site runs no third-party trackers.

Get started

See where your evals and your real users diverge.

Book a scoping call. We'll look at your agent, your current eval stack, and the decision you're trying to make, then come back with a scoped plan and a fixed quote, whether that's a benchmark study, a full human-subject eval, or an ongoing retainer. You leave with a plan, not a deck.

We'll only use this to get in touch about your enquiry. No third-party trackers.

or email hello@oarux.com