BenchMark Research · September 2026 · Krisha Mankarmi \ All rights reserved.
Research question Can we infer what someone is capable of from how they work, not just what they produce?
Professional capability is still judged through proxies such as CVs, interviews, manager reviews and finished work. AI makes those signals weaker because a polished output tells us less about the reasoning, verification or judgment behind it.
BENCH asks whether capability can instead be inferred from behaviour across realistic work.
Capability is a latent state. We cannot observe it directly. We see fragments of it through what someone notices, checks, ignores, prioritises, revises, escalates and decides under changing conditions.
BENCH uses work simulations to generate those observations and combines them into an evolving Professional Twin.
The system needs a structured map of the capabilities it is trying to infer and the relationships between them. This includes reasoning, verification, prioritisation, judgment, communication, adaptability and independence.
The meaning of an action depends on context. BENCH therefore needs to represent the task, actors, information available, goals, constraints and consequences around each decision.
BENCH captures more than the final answer. It records signals from the path through the task, including what someone checks, what they miss, how they revise and how they respond when the situation changes.
The inference engine combines those observations and updates the Professional Twin. The goal is to estimate the capability underneath performance rather than simply score the output.
I do not think professional capability is best represented as a row of independent scores.