SECTION 01Why we publish this.
The AI Exposure Score is the most important number in your Rolespan report. It shapes the curriculum we recommend, the gaps we highlight, and the verdict we open with. If we publish a number that material to your career thinking, we owe you the math behind it.
Most “AI career assessment” products treat their scoring as proprietary magic. We think that's a tell. A score that can't be checked is a score that can't be trusted.
Our principle
If we wouldn't be comfortable defending the methodology to a skeptical reader of the Financial Times, we shouldn't be using it.
This page is written for that skeptical reader. It's longer than a typical product page should be. If you want the short version, the analysis itself contains a “How we got there” summary inside every report.
SECTION 02What the score measures.
The AI Exposure Score is a directional estimateof how much of your role's task mix falls within current AI capability. It runs from 0 to 100, where:
- 0 means none of your day-to-day tasks can plausibly be performed by current AI systems.
- 100 means all of them can.
- 25 to 55 is where most knowledge workers fall in 2026.
The score is a relative readon your task mix versus current capability — not a prediction of when, or whether, you'll lose your job. It's closer in spirit to a credit score than to a weather forecast: a useful proxy for risk that informs decisions, not a literal prediction of events.
What the score is
A frequency-weighted measure of how much of your work falls inside what today's AI can do — computed from your CV's task mix and a per-task estimate of AI capability.
The score is most useful when you read it alongside the task breakdown — high exposure on tasks that produce 60% of your value carries different meaning than high exposure on tasks that produce 10%. The score collapses that into a single number for comparison; the rest of the report explains what's underneath.
SECTION 03How we calculate it.
The calculation has four steps. The math is deliberately simple — we'd rather a method you can audit than a black box.
Step 1 — Extract task categories from your CV
We use a language model to read your CV and extract 12–18 task categoriesthat describe what you actually do. A task category is something like “drafting performance reports” or “negotiating with stakeholders” — not a job title, not a tool name. The same person's CV typically extracts to 14 ± 2 categories.
Step 2 — Weight each task by frequency
Some tasks dominate your time; others are occasional. The model estimates each task's share of your working hours from signals in the CV (how often the task appears, whether it's described as a primary responsibility, role seniority, scope). We then normalize those weights so they sum to 1.0 across all tasks.
Step 3 — Score each task against AI capability
Each task category carries a capability score from 0.0 to 1.0, reflecting how well current AI systems perform that category of work. This is the highest-uncertainty number in the analysis, and it is estimated by the language model rather than looked up from a fixed table. The estimate is informed by:
- Published task-exposure research, primarily Eloundou et al. and the OECD AI & Occupations framework
- Public benchmark performance (HELM, BIG-Bench, MMLU, SWE-bench) as context for what current models can do
- The model's own knowledge of where today's AI is strong and where it still fails
See Data sources for the honest caveat on how much weight to put on these numbers.
Step 4 — Aggregate and adjust
The raw score is the weighted sum of (frequency × capability) across all tasks, then multiplied by 100. We then apply one small adjustment:
- Seniority adjustment (0 to −3 points): more senior roles spend a higher share of their time on judgment, negotiation, and people work, so we nudge their score down slightly — entry roles get 0, executives −3.
We don't currently apply a role-family baseline correction. Calibrating one honestly would need a labeled benchmark we don't have yet, so that term is left at zero rather than guessed.
SECTION 04Data sources.
We want to be straight about where the numbers come from, because the capability score is the part most likely to date the analysis. The per-task frequency and capability values are estimated by a large language model that reads your CV — they are not looked up from a maintained dataset or rated by an expert panel.
That estimate is informed by the public research that frames this field — task-exposure work from Eloundou et al., the OECD's occupational-exposure framework, and public model benchmarks like HELM and SWE-bench. We drew on these to shape how tasks are broken down and what a sensible 0–1 capability range looks like. They are conceptual grounding, listed in the references below; they are not blended in at fixed weights, and no figure on this page is sourced from a private dataset.
The honest implication: treat the capability scores as a well-informed model judgment, not a measured constant. The math on top of them is fully deterministic and reproducible — the same task breakdown always yields the same score — but the inputs carry the uncertainty any estimate does.
SECTION 05A worked example.
Below is a full worked example on Sarah Chen — a Senior Marketing Manager who scores 38 / 100. Sarah is an illustrative profile, not a real user, but the calculation is the real one: every task, every number, shown so you can redo the arithmetic yourself and land on the same 38.
Sarah Chen · Senior Marketing Manager · B2B SaaS · 7 yrs
14 task categories in this example · weighted by frequency · scored against current AI capability| Task category | Freq (fᵢ) | AI cap (cᵢ) | fᵢ · cᵢ |
|---|
| High exposure · AI does this well |
| Drafting blog posts & campaign copy | 0.120 | 0.85 | 0.102 |
| Campaign & creative briefs | 0.080 | 0.80 | 0.064 |
| HubSpot performance reporting | 0.060 | 0.75 | 0.045 |
| Exec updates from dashboards | 0.040 | 0.80 | 0.032 |
| Competitor content research | 0.040 | 0.70 | 0.028 |
| Subtotal · high | 0.34 | — | 0.271 |
| Medium exposure · AI assists, you drive |
| Audience segmentation | 0.060 | 0.55 | 0.033 |
| A/B test design | 0.050 | 0.50 | 0.025 |
| Marketing-mix planning | 0.040 | 0.45 | 0.018 |
| Briefing agencies & freelancers | 0.050 | 0.40 | 0.020 |
| Subtotal · medium | 0.20 | — | 0.096 |
| Low exposure · your moat |
| Cross-functional stakeholder negotiation | 0.150 | 0.10 | 0.015 |
| Brand judgment calls | 0.080 | 0.05 | 0.004 |
| Hiring & developing team of 4 | 0.100 | 0.05 | 0.005 |
| Translating business goals to strategy | 0.080 | 0.10 | 0.008 |
| Crisis & reputation response | 0.050 | 0.05 | 0.003 |
| Subtotal · low | 0.46 | — | 0.035 |
| Raw exposure score | 1.00 | — | 0.402 × 100 = 40.2 |
A few things worth noting about this worked example:
- The low-exposure tasks dominate Sarah's time (0.46 of total frequency) but contribute relatively little to the score because their capability values are low. This is the structural reason senior roles tend to score lower.
- The high-exposure groupcontributes 67% of the raw score from only 34% of her time. This is the textbook “AI will eat the deliverables” pattern.
- If Sarah's role shifted toward execution (more copy, less negotiation), her score could rise 12–18 points without anything changing about her capability — just her task mix.
- The seniority adjustment is small (−2 points here) and deliberately conservative — it ranges from 0 for entry roles to −3 for executives, and there is no separate role-baseline term.
SECTION 06Choosing your top 5 skill gaps.
Skill gaps are not the same as exposure scores. The score asks “how much of your work is AI-exposed?”; skill gaps ask “which specific skills would most reduce your exposure and increase your leverage?”
For each skill we surface, the model scores it 0–1 on three criteria, and we combine them to rank the gaps by marginal impact:
- Exposure reduction. How much would learning this skill lower your effective exposure on your highest-frequency, high-exposure tasks?
- Compounding leverage. Does the skill enable other skills? Briefing AI well, for example, is a foundation skill — once learned, it makes prompt-library work, workflow building, and brand-voice prompting all faster to learn. High-leverage gaps are surfaced first so the list compounds when followed in order.
- Role relevance. How central the skill is to your specific role today, based on the responsibilities and tools the model read from your CV.
Gaps are then ordered highest-marginal-impact-first — a foundation skill that unlocks others is surfaced ahead of a higher-scoring skill that depends on it. The result is a short list that builds on itself.
A deliberate choice
We surface five gaps, not ten or twenty. A short list a user might act on is more valuable than a long list a user will read once and forget.
SECTION 07Matching the curriculum.
Once we know your skill gaps, the model selects courses from the Rolespan catalog that teach them and arranges them into a personal curriculum. It only ever recommends real courses from the catalog — IDs that don't exist are rejected, never invented.
The selection follows a few rules:
- Courses map to your gaps first.The strongest, most direct matches for your top skill gaps lead; where there aren't enough direct fits, role-adjacent courses fill in so the curriculum is substantial rather than thin.
- It's organized into three phases — Start here, Build next, and Go deeper — sequenced as a narrative arc (orient → apply → lead) rather than a flat topic list.
- Each phase is capped and de-duplicated— roughly 2–3 courses to start, 6–7 to build on, 4–5 to go deeper, and no course appears in more than one phase. In total you get about 12–15 courses.
The result is a focused path through the catalog rather than the whole library — enough to act on, ordered so each phase builds on the last.
SECTION 08What the score does not measure.
The Exposure Score is a measure of task-level capability overlap. It is not a measure of:
- Job loss probability.Whether you actually lose your job depends on your company's adoption rate, your manager's preferences, labor market dynamics, and many other factors the score does not see.
- Timing. A high score does not mean the change happens next quarter. Many high-capability tasks remain human-done for years due to trust, regulation, customer preference, or organizational inertia.
- Salary impact. Some roles see compensation rise as AI augments their leverage. Others see it compress. The score does not predict which.
- Your individual employability.A 60-point score for an engineer is very different at a top-quartile performer than at a bottom-quartile one. The score is computed from your CV's task mix, not your skill level within those tasks.
- Industry-specific dynamics. Two marketing managers with identical CVs may face very different real-world AI adoption pressure if one works in regulated healthcare and the other works in growth-stage B2B SaaS.
- The relative quality of human vs. AI output. The capability score reflects whether AI can do the task. Whether it does the task as well as youis a separate question we don't answer here.
SECTION 09Where the methodology can be wrong.
We track six known edge cases where the score is less reliable than usual. If your situation matches one of these, treat the score as a wider range rather than a point estimate. We surface a flag on your report for the first three when we detect them.
- Sparse or vague CVs.CVs that are short or filled with generic phrases (“results-driven leader,” “passionate about innovation”) give the model less to work with, so task extraction is less reliable. We flag this on the report when the extraction comes back low-confidence.
- Hybrid roles. CVs that span two role families (a product-marketing manager, a sales engineer, a designer-developer) are harder to place, so the score can sit a few points either side of its true value. We flag these when we detect them.
- Frontier roles. Roles that involve building AI systems — ML engineers, alignment researchers, AI product managers — are scored with the general framework, but their task mix changes fast enough that scores age quickly. We flag these and suggest re-analyzing periodically.
- Role transitions. If your current role differs materially from your most recent CV entries — for example, you just moved from IC to manager — the score reflects the past, not the present. Re-upload with an updated CV.
- Highly specialized work. Subfields where general AI capability does not predict capability on the specific work — legal practice in a small jurisdiction, niche scientific research, regulated trading roles — are harder to estimate and the model may be over- or under-confident.
- Recent capability shifts. If a major capability shift happened very recently, the estimate may lag. Because capability is estimated per-run rather than stored, re-running your analysis later picks up a fresher read.
SECTION 10Updates & versioning.
Two parts of this method can change over time: the scoring logic (the formula, the seniority adjustment, how skill gaps are ordered) and the capability estimates the model produces as underlying AI gets better at more tasks.
We don't run a fixed quarterly or annual refresh, and we won't pretend otherwise. The scoring logic is versioned — the version is shown at the top of this page — and when we change it in a way that moves scores, we note it in the changelog below. Capability estimates move whenever we update the model and prompt behind the analysis; because they're generated per-run rather than stored in a table, re-running your analysis later is the way to pick up a fresher read.
v1.0 · June 2026
Initial public methodology. The exposure score is computed deterministically from per-task frequency × AI capability, with a small seniority adjustment. Task frequency, capability, and skill-gap scores are estimated by a language model from your CV.
SECTION 11Further reading.
We don't cite live data sources, because the analysis doesn't pull from any — the estimates are made by the model. These are the public works that shaped how we think about task exposure and AI capability, if you want to go deeper on the ideas behind the score.
- Eloundou, T., Manning, S., Mishkin, P., & Rock, D. — "GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models."OPENAI WORKING PAPER · 2023 · CONCEPTUAL GROUNDING FOR TASK-EXPOSURE FRAMING
- OECD. — "Artificial Intelligence and the Labour Market: Occupational Exposure Framework."OECD EMPLOYMENT OUTLOOK · 2024 · CONCEPTUAL GROUNDING
- Stanford CRFM. — "Holistic Evaluation of Language Models (HELM)" & SWE-bench task suites.PUBLIC BENCHMARKS · CONTEXT FOR HOW AI CAPABILITY IS DISCUSSED
- Acemoglu, D. & Restrepo, P. — "Automation and New Tasks: How Technology Displaces and Reinstates Labor."JOURNAL OF ECONOMIC PERSPECTIVES · 2019 · CONTEXTUAL
- U.S. Bureau of Labor Statistics. — O*NET task descriptions, a structural reference for task-category granularity.ONETONLINE.ORG