Case study SOVA · TalentAI Console · 0→1

Nobody trusts
a score they
can't defend.

A hiring platform where AI reads the candidate and a human signs the decision. I designed the layer in between — the one that turns a model output into a judgement someone will stand behind.

Role
Senior Product Designer — end-to-end ownership
Scope
Platform UX, AI insight experience, design system
Partners
Product, Engineering, Data Science, Psychometrics
Decision confidenceReviewing
41%
Not enough evidence to act
Signals3 assessments, 1 work sample, 2 interviews
MissingWork sample not submitted · reviewer 2 of 3 pending
WhyConsistent judgement under time pressure; variation in ambiguous trade-offs

Product The console

One screen that answers: what needs me, and why?

The hiring overview leads with risk and readiness rather than volume. Every number on it is something a person can act on before it becomes a problem.

sova.app / talentai / overview
Last 7 days
All roles
Confidence ≥ 80%
+ New requisition
JP

Hiring overview

Live pipeline health, AI insights and decision readiness.

Pipeline health

3 candidates flagged for incomplete data · 2 insights need human review before shortlisting

Pipeline readiness42↑ 8 this week
Needing review173 high-risk flags
Days to shortlist2.4↓ 18% faster
Insight engagement61%↑ 4% adoption
Candidates awaiting decision3 of 17
AK
Amina KhanSenior Product Designer

Strong analytical reasoning; consistent judgement under pressure

Data completeAI coverage fullHuman review pending
92%
MR
Mason ReidOperations Lead

High process reliability; low ambiguity tolerance

Data partialAI coverage fullWork sample missing
87%
PS
Priya SharmaData Analyst

Evidence still being collected — 1 assessment outstanding

Awaiting assessmentInsight not ready
74%
Candidate insight summaryAmina K.
Data-backed decisions in 8 of 10 scenarios
Stable judgement under time pressure
Variation in ambiguous trade-offs
Insight confidenceHIGH
Evidence sources3 ASSESSMENTS
Last updated2 HOURS AGO
Review candidate
Expand evidence summary
1 2 3 4
01

Risk first, volume second

The banner names what is blocked and why, so nobody has to hunt for the exception.

02

Adoption is a metric

Insight engagement sits next to pipeline health — if people ignore the AI, that is a product failure.

03

State before score

Data completeness, AI coverage and review status are shown beside the number that depends on them.

04

Evidence on demand

The summary reads in five seconds; the reasoning behind it is one deliberate click away.

01 The real problem

It was never a usability problem.

Hiring teams could complete every task in the product and still not know whether they were allowed to believe the result. Three things were eroding confidence at once.

Fragmentation

The journey had no spine

Candidates and hiring teams moved through disconnected steps. Progress was impossible to read, so people re-checked, chased, and stalled instead of moving forward.

Opacity

Intelligence without explanation

The models produced genuinely useful signals. But nobody could see how a result was formed, so the same output was either over-trusted or quietly ignored.

Accountability

Decisions had to be defensible

A rejection can be challenged months later. Completing the workflow wasn't the finish line — producing an outcome a human could justify was.

02 System before screens

I mapped the machine before I drew anything.

Working with engineering and data science, I mapped how assessment data actually moved and where interpretation happened. The map became the team's shared reference — and it showed exactly where confidence was breaking.

L1

Candidate experience

Expectations set up front: what this is, how long it takes, how it's judged.

L2

Assessment & capture

Structured inputs, validated at the point of entry rather than at the end.

L3

AI interpretation

Where the confidence gap lived. Signals became insight — and the reasoning had to surface with it.

L4

Decision confidence

Summary, readiness, and an honest account of what is still missing.

L5

Organisational accountability

Audit trail and traceable outputs that survive a challenge.

Layer 3 is where design earns its keep. Everything above it depends on it being legible.

03 Where I focused first

Three bets, chosen for leverage — not for backlog size.

The roadmap had forty candidates. These three moved the metric that mattered: whether a person felt able to act on what the system told them.

BET 01

Make the first minute legible

  • What this platform is doing for you
  • What happens next, and when
  • Whether you are progressing well

EffectLower drop-off at entry; fewer "where am I?" support tickets.

BET 02

Design for imperfect data

  • Surface gaps early, not at submission
  • Turn errors into a guided next action
  • Always leave a safe recovery path

EffectHigher completion rates and cleaner inputs into the model.

BET 03

Make outputs defensible

  • Plain-language insight summaries
  • Visible confidence and evidence sources
  • Explanation that holds up in an audit

EffectGreater adoption by non-expert reviewers; stronger stakeholder trust.

04 The core shift

From a workflow you complete to a journey you can trust.

Before

Disconnected steps, hidden logic, and a result that arrived without a rationale.

After

One continuous journey that keeps answering the only three questions users actually ask.

Q1 What do I need to do?
Q2 Am I on track?
Q3 Can I trust this outcome?

Product Candidate review

The decision surface.

Where a reviewer forms a judgement they will have to justify. The model's read, the evidence behind it, and what is still missing are held on one screen — with the same logic carried to mobile for reviewers who decide between meetings.

sova.app / talentai / candidates / amina-khan
AK

Amina Khan

Senior Product Designer · Platform team

StageSHORTLIST EVALUATION
Assessment completion94%
Last activityCASE STUDY · 1 DAY AGO
Reviewer status2 OF 3 COMPLETE
Confidence 92%Awaiting shortlist decision
AI candidate snapshotInsight

Demonstrates strong analytical reasoning and consistent judgement across ambiguous decision scenarios. Performance remains stable under time pressure, with moderate variation in risk tolerance during trade-off decisions.

Insight confidenceHIGH
Evidence sources3 ASSESSMENTS
Expand evidence summary
OverviewSignalsWork samplesInterviewsAudit
Interview feedback2 entries
HM
Hiring manager4 days ago

Strong communication and design thinking. Portfolio demonstrates excellent problem framing in ambiguous system work.

TL
Tech lead3 days ago

Comfortable with data constraints. Would probe risk-threshold reasoning further at panel stage.

Recommended reviewer focusAdaptive
Explore decision trade-off reasoning
Validate risk-threshold alignment
Review candidate
Request clarification
9:41
Amina KhanShortlist evaluation · 2 of 3 reviewed AK
92%

Decision confidence

Ready to shortlist
Evidence complete
Why the model thinks this4 signals
Data-backed decisions
Stability under pressure
Ambiguous trade-offs
Structured communication
Probe at panelAdaptive
Decision trade-off reasoning
Risk-threshold alignment
Shortlist
Ask for more
Not now

05 The designed journey

Five stages, each removing one source of doubt.

Every stage pairs what the user experiences with the product decision behind it — because the decision is the part that transfers to the next project.

01

Entry & context

What you're completing, how long it takes, and how it will be evaluated — stated before you start.

Product decision

Set expectations at onboarding rather than explaining them in support articles later.

Removes "where do I start?" hesitation
02

Guided capture

Structured sections that follow how the work is actually thought about, not how the database is shaped.

Product decision

Model the workflow on real decision progression so the interface teaches the process.

Reduces cognitive load mid-assessment
03

Live feedback

Completion and data-quality signals appear while there is still time to fix them.

Product decision

Surface risk while it is recoverable — validation belongs at the point of entry.

Prevents late-stage rework
04

Review & confidence

Readable summaries, a confidence figure, the evidence behind it, and what is still missing.

Product decision

Design the summary for a fast, defensible read — not for completeness of data.

Reviewers can justify the call, not just make it
05

Decision & record

An actionable outcome that carries its own audit trail into the organisation.

Product decision

Structure outputs for accountability so compliance is a by-product of the flow.

Audit-ready by default
Candidate drop-off fell during assessment

Fewer abandoned sessions once progress and expectations were legible

Review decisions got faster and clearer

Shorter time from insight generated to shortlist decision made

Stakeholder confidence rose

Insight features moved from ignored to routinely used in review

Product Before → after

The first build was for analysts. The decision wasn't theirs to make.

Same data, same model, same permissions. The difference is what the screen asks of the person reading it.

Before — Insight explorer

insight explorer — all signals
REASON0.82
STABIL0.77
RISK T0.41
AMBIG0.63
SPEED0.90
VAR0.18
CANDIDATEA1A2A3σIDX
Khan, A.0.820.770.910.060.92
Reid, M.0.740.880.090.87
Sharma, P.0.690.140.74
Osei, D.0.580.610.660.040.63
Lund, E.0.910.850.880.030.89
Bianchi, R.0.470.520.490.020.51

Every signal, equally weighted, with no view on what to do next. Analysts loved it. Hiring managers opened it once.

After — Decision summary

shortlist decision — Amina Khan
Ready to shortlist92%

Strong analytical reasoning and consistent judgement under pressure. Moderate variation in ambiguous risk trade-offs.

Evidence complete3 sources1 reviewer pending
What to probe at panelAdaptive
Decision trade-off reasoning
Risk-threshold alignment
Shortlist
Request clarification
See the evidence behind this

One judgement, its confidence, its gaps, and the two actions that follow. The analysis still exists — it just stopped being the front door.

06 How I made the calls

Four tensions I had to resolve in public.

Designing for AI-assisted judgement meant trading clarity against compliance against speed, with a release date that did not move.

User risk

Explaining the model without drowning the reader

Users either accepted candidate insights uncritically or dismissed them outright. Both failures came from the same gap: no visible reasoning.

More explainabilityCognitive overload
Decision

Lead with a plain summary; let the reasoning open on demand. Depth is available, never mandatory.

Business impact

Built for power users, used by everyone else

The first insight dashboards delighted analysts and overwhelmed the hiring managers who made the actual decisions.

Insight depthWorkflow usability
Decision

Rebuild around decision-ready summaries anchored to real hiring checkpoints, with analysis one level down.

Technical feasibility

An interface that tells the truth about the model

Model confidence and data quality varied case by case. A fixed layout would have implied a certainty the system did not have.

Static clarityModel variability
Decision

Conditional UI states and confidence indicators that adapt to the evidence actually available.

Delivery speed

Shipping transparency in stages

A full explainability framework was the right destination and the wrong first release — it would have delayed learning by months.

Perfect transparencyIncremental validation
Decision

Ship progressive explainability layers that deepen as trust and evidence mature.

Product Progressive explainability

Transparency you can open, not transparency you must read.

Full disclosure up front made reviewers slower and no more accurate. Layering it changed the behaviour: most people read the summary, and the ones challenging a decision opened the evidence — which is exactly the moment it matters.

Try it → expand the reasoning

Decision confidence

Strong analytical reasoning with consistent judgement under pressure.

92%

Data-backed decisions · 8 of 10 scenarios
Judgement stability under time pressure
Consistency across ambiguous trade-offs
Structured communication in work sample

EVIDENCE · 3 ASSESSMENTS, 1 WORK SAMPLE, 2 INTERVIEWS
MODEL v4.2 · LAST SCORED 2 HOURS AGO · LOGGED TO AUDIT TRAIL

07 Delivery

I design with engineering and QA, not for them.

Handoff is where quality usually leaks. I treat delivery as one shared system with five points of contact.

01

Discovery

Co-explore constraints, data structures and regulatory dependencies before committing to a direction.

Kills late technical surprises

02

Definition

Translate flows into system logic, state models and edge cases. Acceptance criteria written together.

Removes ambiguity before build

03

Build

Frequent pairing, implementation checkpoints and live design QA — I don't disappear at handoff.

Keeps intent intact in code

04

Validation

Negative-path testing with QA, error behaviour specs and recovery patterns for the real world.

Resilient, not just functional

05

Scale

Reusable design and logic patterns, component behaviour docs, QA cases aligned to the system.

Growth without redesign

The goal isn't a perfect handoff. It's shared ownership of quality.

Process AI in the loop

How I actually worked: prototype, pressure-test, validate.

The confidence UI had a lot of states — high and low model certainty, missing evidence, partial reviews. Rather than draw every permutation by hand, I built working prototypes with Claude, used it to attack my own logic, and took the survivors to engineering to check they were real.

01 · Prototype

From idea to clickable in an afternoon

I described the confidence card and its states in plain language and had Claude generate a working React prototype. Seeing the low-evidence and conflicting-signal states live — not as static frames — is what exposed which layouts implied more certainty than the model had.

My judgementThe model wrote the code. Deciding which states were honest was the design work.

02 · Pressure-test

Using AI to argue against me

I had Claude role-play a sceptical hiring manager and a compliance reviewer against each screen: where would you over-trust this? What can't you defend? That surfaced the "progressive disclosure" insight faster than a design review would have — full explanations up front made the mock reviewer slower, not surer.

My judgementI chose which objections mattered. AI just made them cheap to generate.

03 · Validate

Feasibility, checked with engineering

Each surviving prototype went to engineering as a runnable artifact, not a spec. Because it already showed the conditional states and data it needed, conversations moved straight to "can the model return this, at this latency?" — turning a week of clarification into a working session.

My judgementThe prototype made the hard questions concrete, so we answered them early.

Working session — confidence states
Me

Build the confidence card for a case where the model is only 58% sure and the work sample is missing. It must not look as trustworthy as the 92% state.

Claude

Done — amber treatment, the score de-emphasised, and an explicit Evidence incomplete row naming what's missing. I also disabled the shortlist action, since acting on it would be the risky path. Want the action re-enabled with a warning instead?

Me

No — keep it disabled. If we can't defend it, the UI shouldn't offer it. Now show me both states side by side so I can check the contrast reads at a glance.

Eng

The 58% path is fine to return live, but "what's missing" needs an extra field from the assessment service. One-day change — worth it. Let's spec that field.

Reconstructed from the working sessions. The point isn't that AI designed it — it's that prototyping in hours let me test more ideas, kill the weak ones earlier, and arrive at engineering with something real to react to.

Foundation Design system

Colour that means something, and states that tell the truth.

In a product where a colour can imply certainty, the palette is a semantic system rather than a style. Indigo is confidence, amber is unresolved, green is verified — and no component may use them for anything else.

Confidence#2B24C8
Confidence · low#E7E5FC
Unresolved#B0730F
Verified#0B7A56
Ink#0E1116
Surface#EDEFF3

State · Complete

Amina Khan92%
Evidence completeAI coverage full
Shortlist

State · Partial

Mason Reid87%
Work sample missingAI coverage full
Request missing evidence

State · Not ready

Priya Sharma74%
1 assessment outstandingInsight not ready
Decision unavailable

One component, three honest states. A score is never shown without the condition of the evidence behind it.

Display / 700Decision confidence
Display / 600Candidates awaiting decision
Body / 400Evidence complete — ready to shortlist
Mono / 500MODEL v4.2 · LOGGED TO AUDIT TRAIL

08 Measurement

How I proved the design was working.

In complex B2B systems, satisfaction scores don't tell you much. I instrumented three questions instead.

Do users trust it enough to act?

  • Task completion confidence
  • Error recovery success rate
  • Time to decision readiness
  • Hesitation and backtracking patterns
  • Qualitative clarity feedback

Is it unlocking real adoption?

  • Workflow completion rate
  • Stage-by-stage drop-off
  • Feature adoption across personas
  • Decision turnaround time
  • Support dependency

Is it reducing delivery risk?

  • Error frequency and recovery
  • Data completeness and validation
  • Engineering rework caused by UX
  • QA defects traced to interaction design
  • AI insight usage and explainability

09 Reflection

What I learned, and what I'd build next.

Learned

  1. Trust in AI comes from progressive transparency, not full disclosure up front.
  2. Decision confidence matters more to users than workflow speed.
  3. System architecture shapes the experience more than visual design does.
  4. Cross-functional alignment early is the cheapest form of rework prevention.

What I'd evolve today

  1. Adaptive explanations that match the reader's expertise.
  2. Confidence scoring that explains how defensible a decision is.
  3. Deeper audit trails built around compliance review workflows.
  4. Reusable insight patterns that generalise across assessment types.