The world builds with AI
it’s time you hire for it

Anyone can ship what the AI wrote. Strata shows you who can catch it when it’s wrong, and proves that’s the signal that predicts who actually performs.

Scroll to learn more
Watch

The two-minute version

What Strata is, why AI judgment is the signal that survives, and how a seeded flaw becomes a defensible hiring decision.

The problem

The world is changing. It's time you hire differently too.

Outdated interviews reward whoever hides AI best, so companies keep hiring people who cracked the test, not people who can do the job.

01

The AI bubble

Every candidate has the same frontier model, and the output all looks brilliant. An interview that grades the output can no longer tell who understood it.

84%build with AI
·
29%trust what it returns
02

Old tests, new cheats

People use AI to crack old-fashioned interviews. The take-home, the LeetCode round, the trivia quiz. Each one is a single prompt away from a perfect answer.

reverse a linked list
03

The pace has flipped

AI ships whole applications in an afternoon. The interview still asks questions written in 2015, so fast-moving candidates get judged by a slow-moving test.

ai · afternooninterview · 2015

Stack Overflow Developer Survey, 2025.

The platform

Not a test you bolt on. The whole funnel.

Companies hire through Strata start to finish: build, invite, assess, decide, sign off. Assessment is the spine; everything around it is ours too.

01

Build

Paste a job description. Strata drafts the workflow: competencies, stages, weights, gates. Edit everything.

02

Invite

Bulk-invite by email or open a public link. Invites, reminders and candidate comms all live in the funnel.

03

Assess

Candidates move through your gated node graph. Every stage produces evidence, not just a number.

04

Decide

Composite + judgment score resolve to GO / NO-GO / REVIEW, each carrying its reasoning trail.

05

Sign off

A human confirms or overrides, logged, auditable end to end. Every override sharpens the system.

Your workflow is a node graph

Six node types, colour-coded. Chain them in any order, weight them, gate them. Tap a node to see what it measures.

gate · top 50%
gate · score ≥ 70
gate · final
optional · drop in anywhere
AI Dev sandboxThe wedge: build with a controlled AI
The wedgeHigh cost tier

A real task in a browser IDE where the only AI is one Strata controls. In sandbox mode its output carries a seeded, plausible flaw. We capture line-level human/AI provenance, AI-reliance %, tool and test runs, then a 12-agent panel scores catch, verify and push-back against your rubric. This is the signal nothing else measures.

weights sum to 1 gates narrow the funnel before the costly nodes run

How the sandbox works

Plant a flaw. Watch what they do with it.

Candidates are never told which task is seeded, or when. Verifying the model’s output is the job, which is exactly what makes it a fair test.

01

Plant

The candidate builds a real task in a browser IDE with an AI assistant Strata controls. In a sandbox task, the assistant’s output carries a seeded flaw, a plausible one. An off-by-one. A swallowed exception. An API that doesn’t exist.

02

Observe

Strata captures provenance, not keystrokes: which lines came from the model, which the candidate wrote, what they ran, what they tested, what they accepted without reading. The flaw either survives to the diff or it doesn’t.

03

Score

A 12-agent panel judges catch, verify and push-back, each agent scoring one dimension, blind to the others, with mandatory cited evidence. A debrief asks the candidate to explain their own submission back.

From evidence to a decision you can defend

Node scores and cited evidence roll into the Context Graph. A weighted composite and the judgment score resolve to one recommendation, never a verdict, always a trail. A human confirms or overrides, and the override is logged.

GO

Clears composite, judgment, must-have coverage and JD-fit.

REVIEW

Borderline on exactly one criterion. A human looks closer.

NO-GO

Falls short, with the reasoning trail that says why.

what the evidence actually looks like

session · candidate buildAI assistant · controlled
14function retryWithBackoff(fn, max) {
15 for (let i = 0; i <= max; i++) {
16 await sleep(2 ** i); // AI: exponential seeded
17 try { return await fn(); } catch {}

Candidate flagged the missing jitter and unbounded first delay before running it, rewrote it, then explained why in the debrief.

Judgment 88Hallucination detectionevidence → ratelimit.ts:16
Pricing

Free to build. Pay when you assess.

Building workflows and designing assessments is free. You pay per candidate-node consumed, priced by tier, so cost tracks the funnel, not a seat you forgot to cancel.

Free

$0forever

Build workflows and design assessments. See exactly what a candidate would.

  • Unlimited workflow building
  • Paste-a-JD workflow architect
  • The Strata template library
  • Rubric editor + Strata defaults
  • Candidate preview
Start free

Pay per assessment

Most popular
Usageper candidate-node

Go live. You only pay when a candidate starts a node, so cost keeps tracking the funnel.

  • Everything in Free
  • Live assessments + invites & comms
  • The AI Judgment Sandbox
  • 12-agent scoring + evidence trail
  • GO / NO-GO / REVIEW decisions
  • Tiered pricing: Low · Med-High · High
Build an assessment

Enterprise

Customannual contract

For teams hiring at volume who need the trust layer, seats and support.

  • Seats, roles & SSO
  • Validity & calibration dashboards
  • Adverse-impact reporting
  • Talent-pool / placement signal
  • Priority support & onboarding
Talk to us

Cost tiers reflect what a node costs to run: Low (Submission, Quiz, Live Interview) · Med-High (Video) · High (AI Dev, AI Interview). Talk to us for volume and design-partner pricing.

Compare

Everyone captures the session. We prove the score.

Legacy tools test AI-solvable syntax. The new wave measures AI-era building but can’t prove it predicts performance. Strata does both, feature-for-feature, then more.

CapabilityStrataAI judgmentLegacy testsHackerRank · CodeSignalAI-era toolsSaffron · Karat
Candidate works with AI (not banned as “cheating”)
Seeded-error judgment scoring
Line-level human vs AI provenance
Context Graph threading every stage
A full multi-node hiring workflow
Adverse-impact analysis built in
Proven validity: the score predicts performance
Zero interviewer hours, fully async

Based on publicly described capabilities as of mid-2026. Session capture and AI transcripts are table stakes now. Seeded-error judgment and proven validity are the differentiators.

What you get

Not a score. A picture of how they engineer.

Every stage produces evidence, and the composite decision is built from it, auditable end to end.

A cited evidence trail, not a number

Every score links to the moment that earned it: the diff line, the transcript timestamp, the test run. You review a case, not a black box. Session replay is built in.

A judgment score that predicts

Catch, verify, push back. Measured, not inferred, and calibrated against real outcomes.

Defensibility, built in

GO / NO-GO / REVIEW with a human override always available and logged. EEOC four-fifths and per-group analysis ship with the platform, on consented data decoupled from scoring.

Zero interviewer hours

Fully async. Candidates run themselves through the funnel; results land in hours, not weeks.

Paste a JD, live in minutes

The architect drafts competencies, stages, weights and gates from the job description. You edit everything. It just skips the blank page. Sign-up to a live assessment in two to three minutes.

A fair candidate experience

Real work instead of trivia, and no surveillance theatre. Integrity signals are advisory, shown with evidence. Nothing auto-rejects a person.

The moat

Reskilll already has the scale and the data.

Capturing a session and scoring it is table stakes. Proving the score predicts who actually performs takes a network no assessment startup can build overnight, and Strata is built on Reskilll’s.

0M+developers in the network
0+hackathons of real build data
Longitudinaloutcome data to calibrate against

Distribution plus longitudinal outcome data is a compounding advantage: candidate acquisition cost near zero, real performance labels to prove the score, and a pool of pre-assessed builders on tap. It’s the one thing a well-funded clone can’t ship next quarter.

Proven, not asserted

Anyone can capture a session and score it. Strata’s validity engine measures Pearson r, AUC and top-quartile lift against real outcomes, and recalibrates GO thresholds where validity actually holds, sliced by role and by judgment dimension.

A warm talent pool

Assessments can draw on builders from the network who’ve opted in, already assessed, already judgment-scored. Supply shows up on day one; you’re reaching people who have demonstrated the exact skill, not buying a cold top-of-funnel.

A compounding loop

Run judgment sandboxes across the base → candidates earn a credential → employers generate inbound → we gather more outcome data → the score gets sharper. Features get cloned. This loop doesn’t.

FAQ

The questions worth asking.

Straight answers on fairness, gaming, and what the score actually means.

Test the judgment, not the typing.

Paste a job description and Strata drafts the workflow. Or start from a template and change everything about it. Your first assessment can be live in minutes.

Applying to a role? Open roles are listed at /jobs, no account needed to browse.