BenchMark Research · September 2026 · Krisha Mankarmi \ All rights reserved.

Research question Can we infer what someone is capable of from how they work, not just what they produce?


The problem

Professional capability is still judged through proxies such as CVs, interviews, manager reviews and finished work. AI makes those signals weaker because a polished output tells us less about the reasoning, verification or judgment behind it.

BENCH asks whether capability can instead be inferred from behaviour across realistic work.

Core hypothesis

Capability is a latent state. We cannot observe it directly. We see fragments of it through what someone notices, checks, ignores, prioritises, revises, escalates and decides under changing conditions.

BENCH uses work simulations to generate those observations and combines them into an evolving Professional Twin.

What BENCH needs to model

Capability ontology

The system needs a structured map of the capabilities it is trying to infer and the relationships between them. This includes reasoning, verification, prioritisation, judgment, communication, adaptability and independence.

Professional world model

The meaning of an action depends on context. BENCH therefore needs to represent the task, actors, information available, goals, constraints and consequences around each decision.

Behavioural evidence

BENCH captures more than the final answer. It records signals from the path through the task, including what someone checks, what they miss, how they revise and how they respond when the situation changes.

Capability inference

The inference engine combines those observations and updates the Professional Twin. The goal is to estimate the capability underneath performance rather than simply score the output.

Capability as topology

I do not think professional capability is best represented as a row of independent scores.