The AIEL Monitor · methods
How the Monitor measures AI demand
The Swedish demand series reads every advertisement on the public job board and asks one question of each: does the text of the role ask the worker to know or do AI? This page documents how that question is operationalised, which words the measure uses, how well it performs against hand-labelled advertisements, and what has changed between versions. Every published figure names the version that produced it, and every version carries a fingerprint, so two numbers built on different definitions cannot be silently mixed.
What we are measuring
Advertised AI-skill demand: does the advertisement, in the text of the role, ask the worker to know or do AI? It is not AI adoption or use by the firm, which requires usage data, and it is not AI exposure, which is a task-based measure of susceptibility. There is one estimand and it is not directly observable: every published series is a bound on it, not the quantity itself. The floor understates it because its lexicon is finite; the broader measure overstates it because it reads text that is not about the role.
We name each series by what the employer does, never by what the job is. The floor is "job ads that ask for AI skills", in the role's own tasks or requirements; the broader measure is "job ads that name an AI skill", anywhere in the text. We avoid the phrase "AI jobs": an advertisement asking for an AI skill is usually an ordinary job that now needs some AI, and conflating the two is the most common source of inflated claims in this literature.
A range, not a point
Floor, ceiling, and the gap between them.
Floor
The advertisement names a specific AI term in the role's own tasks or requirements, and the role builds, integrates or uses AI. This is hard AI-skill demand, and it is what the headline Swedish series reports.
Ceiling
Every advertisement whose role is genuinely AI-related, including those that signal AI only through the surrounding project or field, and those that are about AI rather than doing it. The Lightcast-comparable object is the broader whole-text measure, not this one: the ceiling adds a hand-estimated share of the bare-AI band on top. Note also that the correction reaches only that band and not the whole-text base beneath it, which is 79% of the published figure.
The gap
The AI penumbra: jobs AI touches without requiring the worker to build it. The distance between the bounds is a quantity of interest in its own right, not measurement error.
The lexicon
Where the words come from.
The floor lexicon descends from published, citable sources, and every departure from them is versioned. The core is the deduplicated union of three published keyword lists, which together hold 304 distinct terms. We remove 20 terms that are demonstrably not AI skills (the office and marketing "automation" family, machine code, and four big-data infrastructure tools) and demote four generic robotics terms to a separately reported diagnostic band, leaving a citable core of 280 terms. That band is large in the early years (78% of the broader measure in 2006, 195% in 2007, 16% by 2025) and the demotion therefore steepens the measured trend; the with-and-without series is published alongside the main one. The removals are a validity call, not a definitional one, and each is enumerated in the note's Appendix A.
| Layer | Size | Source |
|---|---|---|
| Published core | 280 terms | Deming and Noray (2020); Alekseeva et al. (2021); Baruffaldi et al. (2020). The same union was used in Engberg et al. (2025), where all three lists are reproduced in full. The general-purpose term "Python" is excluded there and here. |
| Swedish variant layer | 23 concepts, 308 effective core patterns | Ours. No published list covers Swedish, and without this layer every translated concept is missed. Tolerates Swedish compounds (maskininlärningsmodeller and similar). |
| Generative-AI addendum | 75 terms | Ours, anchored to the Lightcast fastest-growing AI skills rather than to our own judgement. Covers post-2022 vocabulary: model families, agentic frameworks, MLOps, vector tooling. |
| Governance cluster | 17 terms | Ours. The EU AI Act, AI ethics, safety and security. Counts towards the ceiling only, never the floor: such postings are about AI rather than asking the worker to build, integrate or use it. |
Admission discipline. New terms enter only through a fixed sequence. An external anchor, a published taxonomy or benchmark, must motivate the candidate; we compute exact impact counts on a full year of advertisements and require near-zero hits in the labelled negative pool; the revised definition must clear a pre-registered gate on the hand-labelled set, with recall up and precision not below the incumbent. Candidates that fail, or that cannot yet be tested, are staged on a public watch list rather than merged.
Version history
Every published figure names the version that produced it.
The third re-freeze. It closes the two lexicon defects v1.2 carried openly, and widens the fingerprint again to cover the pipeline source. The full twenty-year series was reprocessed on it on 7–8 August 2026.
- Plural forms are now matched. Coverage in v1.2 was an accident of which plurals the three source lists happened to duplicate: neural networks matched because the plural is a separate literal entry, while 12 of 24 English multi-word core terms missed theirs. Plurals are generated for every multi-word term plus the named singles llm and chatbot. Multi-word phrases are safe to pluralise mechanically, since a phrase that does not occur in the plural simply never matches; single tokens are not, so they are enumerated instead.
- Bare boosting and bare torch are gone. A parenthetical in the source lists is now split only when it is an acronym: X (machine learning) is a disambiguator, and splitting it had made bare boosting a term matching boosting traffic and boosting sales. Separately, torch is Alekseeva et al.'s entry for the Lua machine-learning framework and matches welding equipment in Swedish advertisements, so it is removed as a corpus false friend. PyTorch is unaffected.
- Both lines move, unlike v1.2 which moved the floor alone. The floor rises from 17,243 to 17,659 advertisements (+2.4%) and the whole-text measure from 37,283 to 38,281 (+2.7%); the raw bare band falls by 136, because an advertisement asking for LLMs was sitting in it purely because the lexicon could not see the inflection. The advertisement count, the adjacent band and the entry-level series move by exactly zero.
- 2006 to 2009 do not move at all, in any measure, which is why the growth figures rise only at their recent end: the whole-text rise since 2006 goes from 21.4-fold to 21.8-fold and the floor from 26.8 to 27.5. This is a change of SHAPE as well as level: 2020 loses 84 whole-text advertisements, the largest single-year fall, while 2022 gains 258, almost all autonomous vehicles.
- The fingerprint now also covers the pipeline source, which holds the deduplication key, the negative-context window and the entry-level regex — all printed as rules the reader is told they can reproduce the number from. A reviewer had changed two of them by mutation with the fingerprint staying green. The pipeline also carries the fingerprint stamp itself, so the hash deliberately excludes those two lines: it covers the rules, not the record of what the rules hash to.
Validation. 88.6% precision [74.0, 95.5] and 37.3% recall [27.7, 48.1] against all AI positives on the 369-advertisement hand-labelled set, with Wilson 95% intervals. The pre-registered gate — recall up, precision not below the incumbent — is met, but both moves sit well inside their intervals, so we claim only that the freeze has not degraded the measure, never that it improved it.
The hand-labelled set is still drawn entirely from 2024 advertisements, so every validation figure describes one year of a twenty-year series. An early-period batch of 120 advertisements from 2006–2013 is drawn across three strata and awaiting labelling; until it is done, the left edge is unvalidated and said to be.
The second re-freeze, correcting a filter fault that suppressed the floor, and widening the fingerprint to cover the matcher source as well as the configuration. The full twenty-year series was reprocessed on it on 5 August 2026.
- The teaching-about-AI filter tested a hard-coded set of four patterns and, when an advertisement had matched an AI term outside that set, fell through to discarding it. A doctoral post requiring PyTorch, TensorFlow and computer vision was thrown out because its benefits section contained the word kurser. The fault suppressed the floor, so every correction to it adds.
- Exactly one measure moved. Summed over all 22 periods the change is zero, not merely small, for the advertisement count, the whole-text measure and its components, the raw bare band, the adjacent band and the entry-level series: the fault sat only on the role-scoped path. The floor rose in every period, from 16,650 to 17,243 advertisements, +3.6%. The identical denominator in each year confirms the fix reclassifies advertisements and never drops one.
- Because the proportional gain is larger in the sparse early years, the floor's own growth from 2006 to 2025 falls slightly, from 53.9-fold to 52.6-fold on counts. The headline whole-text figure, a 21-fold rise, is untouched.
- The fingerprint now covers the configuration minus the hard-negative evaluation list, plus the matcher source. It previously hashed the configuration alone, which was wrong both ways: editing a marker in the matcher changed every published number while the fingerprint stayed identical, and editing the evaluation list forced re-freezes that changed no number.
Validation. 88.2% precision [73.4, 95.3] and 36.1% recall [26.6, 46.9] against all AI positives on the 369-advertisement hand-labelled set, with Wilson 95% intervals. No difference between v1, v1.1 and v1.2 survives its interval, so we claim only that successive freezes have not degraded the measure, never that they improved it.
A known recall defect is carried openly rather than patched: plural forms are missed on many entries: llm and chatbot match in the singular but not the plural, and of 24 English multi-word core terms tested, 12 miss their plural. Measured cost on the 2025 archive is 35 distinct advertisements, about 1.6% of that year's floor. Its direction is the same as the fault above, inflating the bare band and depressing the floor. Fixing it would have changed the fingerprint and invalidated this reprocess, so it was deferred; v1.3 fixes it.
The first re-freeze. Folds three corrections into one release, because all three biased the published series in the same direction and shipping them separately would have moved every number twice within a fortnight, once down and once up.
- The published counting unit becomes the distinct advertisement: headline, employer name and the first 400 characters of the body identical, deduplicated within year. Raw records, v1's unit, are retained in every output file as a robustness line. Repeat postings run from 13.1% of records in 2008 to 48.9% in 2023.
- Four bare product names are date-gated to the date the product came into existence: copilot from 29 June 2021, claude from 14 March 2023, llama from 24 February 2023, gemini from 6 December 2023. Before those dates the tokens match an aviation co-pilot, a given name, an animal and a Swedish company. Multiword forms such as github copilot are unambiguous and ungated, and a gated term is dropped only where the advertisement rests on nothing else.
- Four polysemous terms (prompting, finjustering, bare gpt, bare llm) establish an advertisement on their own only from a stated convention date, and only where no negative-context marker appears within 120 characters. Unlike the product names these words existed throughout and merely acquired an AI sense, so the date is an explicit convention rather than a fact. A strict variant, in which those terms never establish an advertisement alone, is computed in the same pass and published as a robustness bound.
Validation. 90% precision, 34% recall against all AI positives on the 369-advertisement hand-labelled set. The matcher code is byte-identical to v1, so v1.1 is a post-filter over v1's own match sets; every term it touches lies in the generative-AI increment rather than the citable core, leaving core-only results comparable across the freeze.
The gold set is entirely 2024 advertisements, so every product gate is open by construction and the anachronism fix cannot be tested on it. Its evidence is the pre-2009 concentration of the affected advertisements: 301 advertisements between 2006 and 2022 had been classified as AI on no other evidence, a fifth of the 2006 series and under one per cent from 2017.
The first frozen definition. Held core plus the generative-AI addendum, the governance cluster, and three matcher repairs.
- Generative-AI addendum and governance cluster admitted through the admission sequence.
- Matcher repairs: AI/ML folded into the core, multiword plurals, and the glued form generativ-ai.
- Role scoping introduced, restricting the floor to the part of the advertisement describing the role's own tasks and requirements.
- Staged rather than merged: the AI-doing phrase family, the user-tier competence phrases, fleragentssystem, semantic search, edge ai. These remain on the watch list pending validation on a grown gold set.
Validation. 88% precision and 34% floor recall against the 370-advertisement hand-labelled set, against 83% and 4% for the pre-freeze baseline on the same set. The gate was that recall must rise and precision must not fall: met.
Predecessor: held snapshot 23f23bbc32baa0fd.
Public checkability
What is published, and what is not yet.
- This page: the estimand, the definition's composition by layer with its published sources, the full version history with fingerprints, and the validation figures for each freeze.
- Every figure on the Monitor as CSV and SVG, from the footer of the figure itself.
- The underlying advertisements: Platsbanken / JobTech, CC0, so the series can be rebuilt from source by anyone.
- The 2025 upper bound, corrected by hand: 60 advertisements drawn from the bare-AI band and read, of which 12 are genuine AI roles, 20% (95% interval 12–32%). That puts the corrected 2025 ceiling at 1.36% (1.25–1.52) against an uncorrected 2.46%, so the published 2025 range is 0.55% to 1.36%.
- The v1.1 to v1.2 revision log, stating direction and magnitude for every number that moved, and recording that only the floor moved.
- The full technical note (17 pages: estimand, data, lexicon provenance, labelling conventions, validation, benchmarks, the three threats to the trend, and the drift policy). Rewritten 5 August on the corrected figures; two pending diagnostics remain inside it, both named on its own pages.
- The term list itself, as a citable download with its version history.
- The labelling conventions with worked examples. The era-specific correction that turns the raw bare-AI band into the published ceiling is measured for 2025 and provisional, from a classifier, for earlier years; the 2016, 2019 and 2022 samples are drawn but not yet labelled.
- Benchmark levels against Lightcast and the Stanford AI Index. The comparability logic is settled and stated below; the levels themselves are not yet verified against Lightcast's own published material, and we do not publish unverified numbers.
On external benchmarks. Lightcast, which also powers the Stanford AI Index, extracts hand-selected AI skills from the whole posting text with no role scoping, and counts generic "artificial intelligence" as a skill in itself. It therefore corresponds to our broad, whole-text ceiling and not to our role-scoped floor: our floor sits deliberately below Lightcast, and comparability is established at the ceiling. Because the Lightcast series is rising fast, any comparison must be period-matched, and because its AI-occupation band captures only a small fraction of skill-level AI demand, our comparisons target its stacked skill total rather than the occupation band.
Who funds this. The Monitor has no dedicated funder. The grants and institutions behind the research it is built from are listed in full on the support disclosure.
AI-Econ Lab (2026). AIEL Monitor: measuring advertised AI-skill demand, definition v1.3 (fingerprint 8654fe27a3724d06, frozen 7 August 2026). Örebro University and Ratio. Accessed [date].
- Acemoglu, D., Autor, D., Hazell, J. and Restrepo, P. (2022). Artificial Intelligence and Jobs: Evidence from Online Vacancies. Journal of Labor Economics, 40(S1), S293–S340.
- Alekseeva, L., Azar, J., Giné, M., Samila, S. and Taska, B. (2021). The demand for AI skills in the labor market. Labour Economics, 71, 102002.
- Baruffaldi, S., van Beuzekom, B., Dernis, H., Harhoff, D., Rao, N., Rosenfeld, D. and Squicciarini, M. (2020). Identifying and measuring developments in artificial intelligence: Making the impossible possible. OECD Science, Technology and Industry Working Papers 2020/05.
- Deming, D. and Noray, K. (2020). Earnings Dynamics, Changing Job Skills, and STEM Careers. Quarterly Journal of Economics, 135(4), 1965–2005.
- Engberg, E., Hellsten, M., Javed, F., Lodefalk, M., Sabolová, R., Schroeder, S. and Tang, A. (2025). Artificial Intelligence, Hiring and Employment: Job Postings Evidence from Sweden. Applied Economics Letters, early online.