An Article 50 compliance readiness assessment is a per-surface record: for every AI feature you ship into the EU, it captures which paragraph of Article 50 the feature touches, whether you sit on the provider side or the deployer side of that paragraph, what a user actually sees in production today, and where the evidence for that is filed. If the output is one percentage for the whole company, the assessment measured the wrong unit.
Each paragraph of Article 50 is written against a role and a system type rather than against an organisation as a whole. The operative wording sits in Regulation (EU) 2024/1689 on EUR-Lex, with a plain-language summary in the Commission's AI Act Service Desk entry for Article 50. Read against that structure, "we are 72% ready" is not a finding. "Seven of nine surfaces are routed and evidenced; the avatar generator and the alt-text job are not" is.
Judge the assessment on its outputs, not its questionnaire
Most templates you can download are question lists. A question list is an input. The thing worth grading is what sits in the folder afterwards, and there are five artefacts:
- A surface register — one row per AI surface, with an owner's name on it.
- A lane-and-role decision per surface, each with one sentence of reasoning and the document that supports it.
- An exception log holding quoted text, not checkboxes.
- A notice inventory: dated captures of what shipped, from the production build.
- An open-gap list where every line has an owner and a date, not a status colour.
If your assessment produces a score but not these five, it produced a feeling. Everything below is a method for filling the folder.
Step 1: count surfaces, not products
A surface is any place where a user, a customer's customer, or a member of the public meets model output or model-driven interaction. Teams reliably find the chatbot. These are the ones that turn up in the second pass, usually after someone in support mentions them offhand:
1. Support macro suggestions drafted by a model and sent under a human agent's name 2. Subject lines and reply drafts inside an outbound email sequence 3. Alt-text generated for user-uploaded images 4. Meeting or call summaries emailed to external attendees 5. Product descriptions written by a model and synced to a storefront 6. Answer boxes and retrieval snippets sitting above classic search results 7. Per-account onboarding tours and in-app tips 8. Avatar or headshot generation in profile settings 9. The voice bot on the support line 10. Model-written release notes and changelog entries 11. UI strings and help-centre articles produced by machine translation 12. Synthetic sample records shipped inside a demo workspace 13. Sentiment or moderation scores displayed to end users as a label 14. Background replacement, upscaling and "enhance" filters in an uploader
The working rule: if a feature's output leaves your product — email, PDF, published page, exported file, ticket reply, API response — treat it as a surface until the routing step says otherwise. Deleting a row later costs nothing. Discovering it during a customer security review costs a week.
For planning, we budget roughly 20 minutes per surface for intake and 45 minutes for each surface where the role is genuinely contested. A six-surface product lands near two and a half hours for the first pass, and considerably less on re-runs because only changed surfaces get re-scored. That is a planning heuristic from running the exercise, not a benchmark — calibrate it after your first ten surfaces.
Step 2: route each surface to a paragraph and a role
Article 50 splits into four lanes across two roles. The table below summarises who each paragraph is addressed to and where routing goes wrong in practice; the operative wording, including the exceptions attached to each paragraph, sits in the regulation text, so check it there before relying on any summary — this one included.
| Paragraph | Whom the text addresses | Typical product trigger | The routing mistake |
|---|---|---|---|
| 50(1) | Providers of AI systems intended to interact directly with natural persons | Support chatbot, in-app assistant, voice IVR | Counting a footer that says "AI-powered" as the notice |
| 50(2) | Providers of AI systems generating synthetic audio, image, video or text | Image generator, synthetic voice, long-form draft writer | Assuming the upstream vendor's marking survives your export pipeline |
| 50(3) | Deployers of emotion recognition and biometric categorisation systems | Sentiment scoring on calls, engagement analysis on video | Filing it as an HR or analytics tool and never routing it |
| 50(4) | Deployers generating deepfake content, and text published to inform the public on matters of public interest | Synthetic spokesperson video, AI-written news-style posts | Reading "public interest" as "news publishers only" |
One surface can land in two lanes with two different roles. Build a chatbot on a third-party model API and you may sit provider-side for the interaction notice on your own product while the synthetic-output marking question points upstream at the model vendor. The assessment does not settle that; it records what you concluded, in one sentence, with the contract clause, model card or documentation URL behind it. If nobody can name that document in under a minute, mark the lane disputed, put a person's name and a date next to it, and move on. A disputed lane with an owner is a manageable gap. A confidently assigned lane with no source is a landmine. The lane structure and the role split are worked through case by case in Article 50 Transparency Obligations: Four Lanes, Two Roles (2026).
Step 3: score each surface out of 20
Five dimensions, 0 to 4 each. Score the surface, never the company.
| Dimension | 0 points | 2 points | 4 points |
|---|---|---|---|
| Lane assignment | No paragraph named | Paragraph named, no reasoning | Paragraph named, one-sentence reasoning, dated |
| Role assignment | Unassigned or "it depends" | Assigned by assumption | Assigned with the clause, model card or doc URL behind it |
| Notice in production | Nothing user-facing | Exists in staging or in marketing copy | Verified in the production build, mobile and desktop, dated capture filed |
| Exception record | Relying on an exception, nothing written | Exception named | Exception quoted verbatim, dated, with reasoning for this surface |
| Evidence trail | Nothing filed | Filed in an editable doc | Filed append-only with timestamp, actor and surface version |
Read the bands as work states, not verdicts: 0–7 means the surface is unassessed in practice whatever the spreadsheet says; 8–14 means assessed but unevidenced, which is the state most teams are actually in; 15–20 means assessed and evidenced. The number is an internal planning device for sequencing work. It is not a legal position, and no score — from this rubric or any tool, ours included — certifies anything about your position under the regulation.
Report the minimum, not the mean
Fictional example, invented to make the arithmetic visible. Northbeam Analytics runs nine AI surfaces and scores 118 out of a possible 180. The average is 13.1, which reads as "mostly there." Underneath: the chatbot scores 18, six workhorse surfaces sit between 12 and 16, the avatar generator scores 6, and the alt-text job scores 4. The mean is describing the six surfaces that were already fine.
So report two numbers instead of one: the lowest surface score, and the count of surfaces below 8. For Northbeam that is "4, and two." Those two figures survive being read in a hurry by someone who does not work on the product, which is more than can be said for 118/180.
If you are evaluating an off-the-shelf checker rather than building the rubric in-house, apply the same standard to its output — the signals worth grading are set out in this 7-signal rubric for free Article 50 checks.
Step 4: exceptions get quoted, not summarised
Article 50 carries exceptions, and this is where readiness folders quietly rot. A row that reads `Exception: obvious` is worth nothing six months later, because it does not say obvious to whom, on which screen, at which point in the flow, or who decided.
Give every exception three fields: the quoted limb of the text you are relying on, the date you read it, and one sentence of applied reasoning about this specific surface. Quoting rather than paraphrasing takes fifteen seconds and removes the risk that a summary drifted from the wording between the first pass and the review.
Whether a given exception covers a given surface is a legal question, not a product one. Where the answer is genuinely contested — and for edge cases in the creative, assistive and public-interest wording it often is — the honest output of the assessment is a clean, well-framed question for counsel, not a manufactured yes.
Step 5: dates you verify rather than assume
Article 50's transparency provisions apply from 2 August 2026, and a limited transition may extend provider-side machine-readable marking to 2 December 2026 for eligible systems; both dates and the eligibility scope are worth checking against the current AI Act Service Desk page rather than against a blog post, this one included, before you plan a release train around them. The timeline and what each date changes for a product team are laid out in AI Disclosure Compliance Deadline 2026.
Put the date you last checked into the register. A dated verification is an artefact. An undated belief is not.
The shelf life nobody plans for
Here is the failure that no template addresses: a readiness assessment describes a product build, and product builds change weekly. An assessment completed in March against a February build is a photograph, and by June it documents something you no longer ship.
Version the register against change triggers rather than the calendar. Re-score a surface when any of these happen:
1. The model vendor or model family changes 2. A system-prompt change alters what the output claims to be 3. The surface's copy or first-run flow changes 4. A new EU locale or market goes live 5. Human-in-the-loop review is removed and sending becomes automatic 6. An existing feature gains a new output channel — email, PDF, public page, API 7. Vendor terms or model-card language about output marking changes 8. Any user asks whether a human wrote it
Trigger 8 is the cheapest signal you will ever get and the one most often thrown away in a support queue. Route those tickets to whoever owns the register.
When a trigger fires, re-score that surface only. The register version bumps; the other rows keep their dates. A naming convention that survives audit-by-stranger: `A50-<surface-slug>-<YYYYMMDD>-v<n>`, so `A50-alt-text-20260824-v3` reads correctly to a person who has never seen your folder. Add a calendar floor of twelve months for surfaces that never trigger, so dormant features do not silently age out.
What typically goes wrong
The assessment was run at the wrong altitude. Someone answers 40 questions on behalf of the company, produces a maturity score, and no individual feature is ever routed. Symptom: the output contains no surface names.
The vendor assumption. "Our model provider handles marking" appears in the register with no clause behind it. Nobody has checked whether the marking survives the surface's own export, screenshot, PDF or re-encode step.
Staging drift. The notice was verified in a preview build. Production ships behind a feature flag that renders it only for accounts created after a certain date.
The 768px notice. The disclosure sits in a sidebar that collapses on mobile. Check every capture at 360px and with a screen reader before it counts as evidence.
Editable evidence. The folder is a shared doc anyone can revise. Once a record can be silently rewritten, its value as evidence drops to roughly zero, which is the entire argument for append-only storage.
One heroic owner. A single person holds the register in their head and their spreadsheet. They change teams. That is the most common way an assessment stops existing.
What an evidence file looks like when it holds up
| Artefact | Minimum content | Where it breaks |
|---|---|---|
| Surface register | ID, owner, model and vendor, output channels, EU exposure, last scored date | Lives in one person's drive and leaves with them |
| Lane-and-role decision | Paragraph, role, one-sentence reasoning, supporting document | Reasoning reads "obvious", with no addressee named |
| Notice capture | Dated screenshot or recording, build ID, viewport, locale | Captured on staging, or only at desktop width |
| Exception record | Quoted limb, date read, applied reasoning, reviewer | A checkbox where a quote belongs |
| Change log | What changed, when, who re-scored, resulting version ID | Product ships; the register is never versioned |
| Verification run | Timestamp, surface or URL, what was checked, result | Run once at launch and never again |
One boundary worth stating plainly, because it gets blurred in vendor material: machine-readable marking techniques — HTML attributes, JSON-LD blocks, response headers — are advisory metadata. They are not signed provenance and not C2PA certification, they survive a screenshot poorly, and they are stripped by plenty of ordinary pipelines. Record them as what they are. DiscloseKit's own workflow is a deterministic checker, a disclosure widget, a verification step and an append-only evidence log; nothing in that shape, ours included, can tell you your legal position.
Twelve questions you can answer in one sitting
1. Can you list every AI surface exposed to EU users without opening the codebase? Pass if engineering adds nothing new when you read the list aloud to them. 2. Is the Article 50 paragraph named in writing for each surface? Pass if you recorded paragraph numbers, not category labels like "generative". 3. Is the role recorded with a supporting document? Pass if you can open that clause or model card inside a minute. 4. Does any surface carry a disputed role with no owner and no date? Pass if the count is zero. 5. Was the user-facing notice checked in the production build this quarter? Pass if there is a dated capture with a build ID. 6. Does the notice survive 360px, a screen reader and a locale switch? Pass if you have captures for all three. 7. Where you rely on an exception, is the limb quoted verbatim? Pass if no exception row is shorter than a sentence. 8. Do you know which outputs leave through channels the notice does not follow — email, PDF, API, export? Pass if the register has an output-channel column. 9. Can you show what a surface looked like three releases ago? Pass if captures are versioned, not overwritten. 10. Is the evidence store append-only? Pass if nobody on the team has permission to edit a filed record. 11. Does a named person own re-scoring after a model swap? Pass if you can say the name without checking. 12. If a customer's security team asked for the file tomorrow, how long to assemble it? Pass if under a working day, without emailing a vendor.
Eight or fewer passes and the useful next step is Step 1 again, not a tooling decision.
Three questions that come up every time
Does using a third-party model API move this off our plate? It changes what evidence you can produce, not whether the surface needs routing. The vendor's documentation becomes an input to your record — the clause you cite in the role field — rather than a replacement for it. Where a lane genuinely straddles you and the vendor, the assessment's job is to state the split and the reasoning; settling a contested split is a question for counsel.
Is this the same as a conformity assessment? No — they are different instruments within Regulation (EU) 2024/1689, and the conformity assessment procedures attach to systems classified as high risk, while the Article 50 transparency provisions run on their own track. A surface can be outside the high-risk classification and still sit squarely inside an Article 50 lane. Check which parts of the text apply to your systems against the regulation itself.
How often do we re-run it? On the eight triggers above, with a twelve-month floor for surfaces that never fire one. Quarterly full re-runs mostly generate work for surfaces that did not change.
The question for your next product review
Which surfaces shipped or changed since the last scoring round, and whose name is next to each of them? A readiness assessment that is not versioned against your release train is an accurate description of a product you no longer run — and it will read as thorough right up until someone opens it next to the live build.
