An AI Act compliance tool comparison only becomes useful once you stop comparing brands and start comparing categories, because the five categories on this market produce five different artefacts — and four of them never touch the sentence your user actually reads on screen.
The ranked listicles blur that. They put a free risk classifier, a GRC automation suite, an AI governance platform, a runtime observability layer and a disclosure layer into one table and score them on shared checkboxes like "audit trail" and "policy templates". A team can buy the top item from one of those lists, complete every onboarding step, and still ship a support chat widget with nothing in the interface that tells anyone it is a chat with an AI system.
Article 50's transparency provisions have applied since 2 August 2026, and a limited transition may extend provider-side machine-readable marking to 2 December 2026 for eligible systems. Check both dates and their scope against the European Commission's official AI Act Compliance Checker rather than against a vendor's marketing page, and treat the question of whether your particular system falls inside a given lane as one for qualified counsel. This article is not legal advice; it is a purchasing method.
Five categories, five different artefacts
| Category | What it produces | Who runs it week to week | What it hands back to you |
|---|---|---|---|
| Free risk classifier | A routing verdict: which tiers and lanes may apply | Anyone, in one sitting | Everything downstream of the verdict |
| GRC automation suite | Control library, policy set, auditor workspace | Compliance or security | Product surface work, notice copy, marking |
| AI governance platform | System and model inventory, risk assessments, framework mapping | Governance or risk lead | The rendered UI, live verification |
| Runtime observability and guardrails | Traces, filters, incident records, prompt and output logs | Platform engineering | Classification, notice text, dated evidence |
| Disclosure layer | The user-facing notice, output marking, live verification, evidence entries | Product engineering | Risk-tier work beyond transparency |
DiscloseKit, which publishes this guide, sits in the fifth row: a deterministic checker with no model in the compliance core, a disclosure widget, a verification step and an append-only evidence log, hosted in the EU. Treat that row's description as self-interested, and score it with the same six axes below as everything else. A vendor that flinches at being scored on exit cost is telling you something.
Two free classifiers are worth running before you talk to anyone: the Commission's checker linked above, and the independent EU AI Act Compliance Checker maintained alongside the annotated Act text. They cost nothing and they narrow the conversation from "AI compliance" to a specific set of lanes. They also stop where every free classifier stops: at the verdict.
The question that sorts vendors faster than a feature matrix
Ask one thing on the first call. *After your tool finishes its job on our checkout assistant, what artefact exists that did not exist before, and where does it live?*
If the answer is a report, a risk register entry or a dashboard tile, you are buying analysis. If the answer is a rendered string inside the product plus a dated record of that string being live at a URL, you are buying execution. Both are legitimate purchases. Buying analysis while believing you bought execution is the failure mode this entire market produces, and it usually surfaces the first time someone outside the team asks to see what a user saw on a specific date.
The follow-up question is sharper: *which parts of that answer are in the contract, and which are in the docs?* Docs change on a deploy. Contracts do not.
A 30-cell coverage matrix you can fill from vendor docs
Six duty cells, five categories. The four lanes below follow the structure of Article 50 — an interaction notice, machine-readable marking of synthetic output, a notice for emotion recognition and biometric categorisation, and disclosure for deepfakes and AI-generated public-interest text — with the paragraph split and the provider/deployer allocation being exactly the part to confirm against the official source for your own facts.
| Duty cell | Free classifier | GRC suite | Governance platform | Runtime layer | Disclosure layer |
|---|---|---|---|---|---|
| Lane 1 — interaction notice in the UI | Routes only | Policy only | Policy only | Partial | Renders it |
| Lane 2 — machine-readable marking of output | Routes only | Not covered | Policy only | Partial | Emits it |
| Lane 3 — emotion recognition / biometric notice | Routes only | Policy only | Policy only | Not covered | Renders it |
| Lane 4 — deepfake and public-interest text disclosure | Routes only | Policy only | Partial | Not covered | Renders it |
| Dated evidence retention | Not covered | Strong | Strong | Partial | Strong |
| Change detection when a surface ships | Not covered | Not covered | Partial | Strong | Partial |
That grid is a category-level sketch, not a vendor scorecard. Two products in the same column can differ sharply, and several vendors sell across two columns. Print it, then fill one copy per shortlisted vendor using only their own public documentation, and leave every cell you cannot source blank. An empty cell is a finding. A cell filled from a sales call is not evidence of anything.
The pattern that emerges from most shortlists: the paperwork rows are well served and the top four rows are thin. That asymmetry exists because inventory and policy tooling generalises across regulations, while a notice in a specific React component does not.
Score six axes out of four, with two hard floors
Twenty-four points, six dimensions, no weighting games. Anchors at 0, 2 and 4 so two people scoring separately land in the same place.
1. Surface reach. 0 — produces documents only. 2 — provides copy and snippets you paste in yourself. 4 — renders in the product and can be verified from outside it.
2. Role handling. 0 — asks you to pick provider or deployer once, globally. 2 — per system. 4 — per surface, and can hold both roles on the same feature when you build on a third-party model API and also deploy the result.
3. Verdict determinism. 0 — an opaque score. 2 — a rules engine you cannot inspect. 4 — same inputs give the same verdict, the verdict names the paragraph it relies on, and the rule set carries a version number.
4. Evidence quality. 0 — a PDF export you can regenerate at will. 2 — timestamped entries that an admin can edit. 4 — append-only entries containing the verbatim rendered artefact, the verification method, and a supersedes pointer.
5. Change resilience. 0 — vendor rule updates silently reclassify your past records. 2 — updates are announced. 4 — old verdicts stay pinned to the rule version that produced them, and new verdicts appear as new records.
6. Exit cost. 0 — screenshots of a dashboard. 2 — CSV of metadata. 4 — a full export you can read and re-verify without the vendor's software.
Two hard floors: a zero on evidence quality or on exit cost disqualifies the tool as the system of record for that lane, whatever the total says. Below 14 out of 24 in this framework, plan on supplementing it. Between 14 and 19, it can carry one or two lanes. At 20 or above, it can be the record for the lanes it covers, and no further. If you are evaluating a single product rather than a field, our longer walkthrough of an EU AI Act Article 50 compliance tool covers the intake fields and the control matrix in more depth.
The two-hour bake-off: six cases, twenty minutes each
Do this with a sample of twelve real surfaces from your product, not with the vendor's demo data. Twelve is enough to expose drift and small enough to finish in a morning. If you have not built the surface list yet, start with the inventory pass in our EU AI Act compliance checklist — the bake-off is worthless without it.
Case 1 — the obvious chatbot. A support widget labelled "AI Assistant" in 24px type, with an avatar and a disclaimer above the input. Does the tool let a human record and justify an exception, or does it force a notice and call the job done? The right behaviour is a recorded judgement with a reason field, not a silent pass in either direction.
Case 2 — someone else's model. Your summariser calls a third-party API. Which role does the tool assign you, and does it explain the reasoning? Watch for tools that assign one role to the whole account. If the tool cannot express "provider for the surface we built, deployer for the model we consume", it will misfile most of a modern SaaS product.
Case 3 — the retouched marketing image. A product photo with an AI-extended background, published on your pricing page. Does the tool route it into the deepfake lane, out of it, or into a "needs human judgement" queue? All three answers can be defensible. An answer with no reasoning attached is not.
Case 4 — the AI-drafted incident update. Text generated for your public status page about a service outage. This is the case most tools fumble, because it lives in the public-interest text lane rather than the chatbot lane, and it is published by marketing or SRE rather than product.
Case 5 — the rename. Change "Smart Reply" to "Compose" and redeploy. Does the record version, or does it overwrite? Overwriting destroys your ability to answer questions about the past, which is the only thing the record is for.
Case 6 — the auditor's question. "Show me that the notice was live on 3 March at 14:00." Time how long it takes. Under two minutes with a single export is a working system. Twenty minutes of Slack archaeology means the evidence lives in people, not in the tool.
Score each case pass, partial or fail before you move on. A vendor that passes cases 1 through 4 and fails 5 and 6 has built a good classifier and no memory.
Where machine-readable marking claims quietly break
Marking is where comparison articles get loosest, because "machine-readable marking" sounds like proof and behaves like a hint. HTML attributes, JSON-LD blocks and HTTP response headers are advisory metadata. They are not signed provenance, they are not C2PA content credentials, and they do not survive a user pressing Cmd+Shift+4.
Run each shortlisted vendor's marking channels through five transforms and make them tell you, in writing, which cells they claim:
| Marking channel | Screenshot | Social re-upload | PDF export | Copy-paste of text | CDN resize / re-encode |
|---|---|---|---|---|---|
| HTML attribute on the element | Lost | Lost | Lost | Lost | Not applicable |
| JSON-LD block in the page | Lost | Lost | Lost | Lost | Not applicable |
| HTTP response header | Lost | Lost | Lost | Lost | Not applicable |
| Embedded file metadata (EXIF / XMP) | Lost | Commonly stripped — test yours | Depends on the exporter | Not applicable | Depends on the pipeline |
| Visible on-screen or in-frame label | Survives | Survives | Survives | Lost | Survives |
Fifteen cells, and only one row holds up across the board — the one that is visible to a human. That is not an argument against the machine-readable channels; the text asks for detectability by machines, and a visible label alone does not deliver that. It is an argument for running both, and for being precise in your own documentation about what each one demonstrates. A vendor who describes an HTML attribute as "tamper-proof" or "provenance" has either misunderstood their own product or is hoping you will. Our scorecard for AI content labeling tools goes through the channel-by-channel tests in detail.
The 11 fields that turn a note into evidence
Whatever the tool calls its records, the per-surface file wants these:
1. Surface identifier and the human name your team actually uses for it 2. The role you claimed for that surface, with a one-sentence reason 3. Lanes in scope, and lanes explicitly ruled out 4. Rule-set version and the date the verdict was produced 5. The rendered notice text, verbatim, not a template reference 6. Placement: route, trigger, and position in the flow 7. Timestamp of the live verification plus the method used (HTTP fetch, DOM check, screenshot hash) 8. Approver — a named person, not a team alias 9. Linked release or commit 10. Next scheduled re-verification date 11. Supersedes and superseded-by pointers
Field 11 is the one people skip and the one that matters when someone asks about last March. And the test for the whole set is blunt: if a field can be edited later without leaving a trace, it isn't evidence — it's a note.
Build or buy: count evidence events, not seats
The seat-count question hides the real driver, which is how often a record has to be created or refreshed. Count it directly:
> evidence events per year = surfaces × lanes in scope per surface × (scheduled re-verifications per year + releases per year that touch the surface)
A worked example, invented for this article: Northbay Analytics, a fictional 40-person SaaS with nine AI surfaces. Seven surfaces sit in one lane, two sit in two lanes, giving 11 surface-lane pairs — call it 14 once you include the three surfaces that also generate text for a public-facing page. Quarterly re-verification is four events; each surface is touched by roughly three releases a year. That is 14 × 7 = 98 evidence events a year, or roughly two a week.
Thresholds worth arguing about, in this framework and adjustable to your release cadence:
- Under 40 events a year. A spreadsheet, a folder of dated screenshots and a calendar reminder hold up. Buying a platform for this buys process, not capability.
- 40 to 150. Manual upkeep starts lagging behind releases. The minimum viable purchase is something that verifies the live surface automatically and writes a dated record without a human remembering to.
- Above 150. The bottleneck moves from verification to reconciliation — knowing which of 300 records are current. This is where inventory-first platforms start earning their price, and where a disclosure layer alone stops being sufficient.
Note what does not appear in that formula: company headcount, funding stage, or whether anyone has said the word "governance" in a board meeting.
What teams get wrong when they choose
Buying the classifier and calling the surface covered. The free checkers are good and they are honest about their scope. The verdict is the beginning of the work. Teams routinely file the PDF and move on.
Letting the wrong function own the inventory. When the AI surface list lives with compliance, it goes stale within two release cycles, because compliance is not in the pull requests. When it lives with product and compliance reads it, it stays roughly current. The tool cannot fix this, and no vendor will tell you it is your org chart rather than their software.
Screenshots without a chain. A folder of PNGs proves that someone took screenshots. Without a hash, a timestamp from outside the file's own metadata, and a link to the release, it establishes very little about what a user saw.
Silent reclassification. A vendor updates its rule set, your dashboard turns green, and nobody can reconstruct what the verdict was in April. Ask specifically how historical verdicts are pinned. This is axis five, and it is the axis that most demos skip.
Reusing the GDPR record. Records of processing and Article 50 transparency records answer different questions and have different owners. Copying one into the other produces a document that satisfies neither reader.
Three fictional stacks, scored
All three are invented profiles, not customers.
Osprey (seed stage, three AI surfaces, one role). One chat assistant, one AI-drafted email helper, one summariser on someone else's API. Roughly 3 × 1 × 6 = 18 evidence events a year. A free classifier plus a disclosure layer scored 21/24 covers the ground; a governance platform would score well and solve a problem Osprey does not have yet. Total spend on process should stay below one engineer-day a quarter.
Meridian (Series B, 14 surfaces, both roles, no high-risk classification). Around 98 events a year, the Northbay shape. The interesting decision is not which category to buy but which two: the coverage matrix rows they cannot fill are lanes 1, 2 and 4 plus change detection. A disclosure layer at 20+ on the scorecard combined with the change detection they already have in CI beats a single platform scoring 17 across everything.
Halden (enterprise, 60+ surfaces, some routed into higher-risk tiers by the official checker). Above 400 events a year, plus obligations well outside transparency. Here an inventory-first governance platform is doing the load-bearing work, and the question becomes whether it reaches the rendered UI — axis one — or whether that stays a separate layer. In the profiles we have modelled, it stays separate.
Nine things to get from the vendor in writing
1. Which of the six duty cells does your product produce an artefact for, and which does it only document? 2. How do you assign provider and deployer roles when both apply to one feature? 3. What is your rule-set versioning model, and are historical verdicts pinned to the version that produced them? 4. Which marking channels do you emit, and which of the five transforms above does each survive? 5. Can an administrator edit or delete an evidence entry, and does that leave a trace? 6. What exactly is in a full export, and can we re-verify it without your software? 7. Where is the data hosted, and who is the processor? 8. How does the tool learn that a surface changed — polling, webhook, CI hook, or a human? 9. Which of the answers above are contractual, and which are documentation subject to change?
Question 9 is the one that changes the conversation. Ask it last.
Questions people ask before they buy
Is the official Compliance Checker enough on its own?
It answers the routing question — which rules may apply to your system — and it is the right first stop precisely because it is free and official. It does not render a notice in your product, mark an output, verify that anything is live, or keep a dated record. Whether you need anything beyond the verdict depends entirely on what the verdict says for your surfaces.
We already run a GRC platform. Do we need a second tool?
Fill the 30-cell matrix for your existing platform first. If the bottom two rows are strong and the top four are empty, you have a coverage gap in the product surface rather than a tooling gap in compliance, and the cheapest fix is often engineering work plus a verification step, not a second subscription. If the top four rows are partly filled, run the six bake-off cases against the platform you already pay for before you shortlist anything new.
Does buying a tool mean we are covered?
No tool settles that. A checker produces a routing verdict from the inputs you gave it. A widget renders a string. A log records that the string was live at a moment in time. Whether those artefacts add up to meeting the duties that apply to your specific system, in your specific role, is a legal judgement about your facts — and that belongs with qualified counsel, informed by the records the tooling produced.
The case this framework doesn't settle
Every scoring method above assumes you can name a surface and assign it a role. The awkward case is the feature where you are plainly both: you built the interface and you consume a foundation model through an API, so one half of the sentence a user reads is yours to design and the other half depends on marking emitted upstream by a party you cannot audit. No tool in any of the five categories resolves that split for you. What a tool can do is record which half you claimed, on what reasoning, and on what date — which is what you will want in hand when someone asks.
Start by running the six bake-off cases against whatever you already own. Most teams discover their gap is three rows wide, not five, and that changes the shortlist before the first sales call.
