arXiv is now an independent nonprofit! Learn more
License: CC BY-NC-ND 4.0
arXiv:2609.39537v1 [cs.CY] 30 Sep 2026

A Reusable Semantic Web Framework for Evidence-Grounded Fundamental Rights Impact Assessments under the EU AI Act

Conference: Fifth European Conference on Algorithmic Fairness; September 02–September 04, 2026; Ghent, BE
Faith Olopade Affiliation: Trinity College Dublin, Dublin, Ireland email: olopadef@tcd.ie , Delaram Golpayegani Affiliation: ADAPT Centre, Trinity College Dublin, Dublin, Ireland email: golpayes@tcd.ie and David Lewis Affiliation: ADAPT Centre, Trinity College Dublin, Dublin, Ireland email: dave.lewis@tcd.ie
2026
Abstract.

The EU AI Act (Art. 27) requires deployers of high-risk AI systems to conduct Fundamental Rights Impact Assessments (FRIAs) before deployment, yet the evidence needed for credible assessments is fragmented across incompatible incident repositories, risk vocabularies, and legal texts. We present a reusable Semantic Web-based framework that consolidates this evidence for two high-risk public sector categories: employment and worker management (Annex III(4)) and access to essential public services (Annex III(5)(a)). A curated 150-record corpus is annotated along four axes using keyword, LLM, and hybrid methods and serialised as a SPARQL-queryable knowledge graph of 1,351 RDF triples. Five FRIA demonstration scenarios surface 103 records (68.7% coverage). Evaluation against a 69-record gold standard reveals that LLM-assisted classification of the employment domain achieves only κ=0.045\kappa=0.045, a cautionary result for automated fairness-related evidence retrieval in this domain. All artefacts are released openly to support adoption by regulators, national authorities, and SMEs.

Keywords:
EU AI Act, Fundamental Rights Impact Assessment, knowledge graph, Semantic Web, LLM classification, algorithmic fairness, public sector AI

1. Introduction

Government agencies across Europe are deploying large language models (LLMs) at scale: welfare eligibility triage, citizen-facing chatbots, recruitment screening pipelines, and caseworker decision support (Bommasani and others, 2021). The appeal is straightforward, but the risks to fundamental rights are direct. Bias amplification, hallucinated outputs, and opacity bear on individuals’ access to work and essential services in ways that are not always recoverable (Weidinger and others, 2021; Bender et al., 2021).

Real-world cases demonstrate what inadequate assessment looks like in practice. In the Netherlands, the Tax Administration’s algorithmic risk scoring system embedded racial profiling and operated without meaningful oversight, wrongly accusing tens of thousands of families of benefit fraud (Amnesty International, 2021). The SyRI case established that automated profiling in public welfare contexts can constitute a violation of the European Convention on Human Rights (District Court of The Hague, 2020; Wieringa and others, 2023). In employment, the Mobley v. Workday litigation alleges that an AI-powered résumé screening tool systematically discriminates on the basis of age, race, and disability. A persistent pattern runs through these incidents: high-stakes public sector AI can cause large-scale rights violations when risk assessment is inadequate.

The EU AI Act responds to this directly. Article 27 requires that deployers of public sector high-risk AI systems, listed in Annex III, conduct a Fundamental Rights Impact Assessment (FRIA) before deployment, documenting how the system may affect rights including non-discrimination, privacy, and good administration (European Parliament and Council of the European Union, 2024). The European Union Agency for Fundamental Rights has noted that many deployments still lack the tools and evidence base to conduct meaningful assessments (European Union Agency for Fundamental Rights, 2025).

The evidence deployers need is not available in a usable form: it is fragmented across sources, described inconsistently, and not interoperable, so assembling it for a given assessment demands substantial manual effort. Incident data resides in repositories such as AIAAIC (AI, Algorithmic and Automation Incidents and Controversies) and AIID (AI Incident Database), which use free-text narratives and community-developed tags rather than AI Act terminology (Pownall, 2024; McGregor, 2021). Existing Semantic Web risk vocabularies—such as Data Privacy Vocabulary (DPV) (Pandit et al., 2024), Vocabulary of AI Risks (VAIR) (Golpayegani et al., 2024b), and AI Risk Ontology (AIRO) (Golpayegani et al., 2022)—provide structured, machine-readable taxonomies aligned with the regulatory framework but are not integrated into FRIA workflows (W3C Data Privacy Vocabulary Community Group, 2022; Golpayegani et al., 2024b). Given that legal obligations of the AI Act, including FRIA obligations, and the EU Charter of Fundamental Rights are expressed in natural language, they do not support automation (European Parliament, Council of the European Union, and European Commission, 2000). Therefore, deployers must manually reconcile these disparate sources with no reusable infrastructure to connect incident evidence, risk vocabularies, and regulatory obligations.

Bridging that gap, this paper presents an interoperable framework for AI Act’s FRIAs using Semantic Web technologies, with a focus on two categories of Annex III high-risk AI applications. The full technical implementation is described in the accompanying dissertation (Olopade, 2026). This extended abstract presents the framework, its key empirical findings, and their implications for algorithmic fairness research and practice.

2. Framework

2.1. Corpus Construction

The corpus combines three complementary evidence sources: approximately 100 records from the AIAAIC incident repository (Pownall, 2024), 30 entries from the U.S. Federal AI Use Case Inventory (Office of Management and Budget, 2024), and 20 European Court of Human Rights cases from HUDOC (Council of Europe, 2024). Together these provide 150 records covering LLM-related and analogous AI deployments in employment and essential public services. The three sources triangulate across evidence types: AIAAIC captures real-world failures; the U.S. Federal Inventory documents active government deployments; ECtHR cases provide authoritative legal analysis of rights violations.

Each record is classified along four axes: (1) Annex III high-risk AI domain (employment or essential services); (2) implicated EU Charter rights (multi-label); (3) risk pattern (bias/discrimination, privacy breach, procedural unfairness, lack of transparency, or other); and (4) causal factor (data quality, model design, deployment context, or oversight failure).

2.2. Semantic Schema and Knowledge Graph

The schema is grounded in Semantic Web standards and aligned with DPV (W3C Data Privacy Vocabulary Community Group, 2022; Pandit et al., 2024), VAIR (Golpayegani et al., 2024b), AIRO (Golpayegani et al., 2022), and the FRIA ontology (Rintamäki and Pandit, 2025). Reusing established vocabularies reduces the learning burden for adopters and maximises interoperability with existing compliance tooling. The annotated corpus is serialised in two formats: a Turtle RDF knowledge graph of 1,351 triples (193 nodes, 965 edges) enabling information retrieval using the SPARQL query language11 1 https://www.w3.org/TR/sparql11-query/, and 150 JSON-LD records for web-compatible consumption.22 2 All schema artefacts, the knowledge graph, and the SPARQL queries are available at https://github.com/faitholopade/Dissertation.

2.3. Annotation Pipeline

Three annotation methods are implemented and compared. A keyword baseline uses synonym-expanded dictionary matching against controlled vocabularies. An LLM classifier applies Claude Sonnet (Anthropic, 2024) with zero-temperature sampling, few-shot prompting, and structured output constraints. A hybrid method resolves conflicts between keyword and LLM outputs using priority logic, retaining keyword classifications for domain and LLM classifications for rights where each method performs better.

2.4. Regulatory Crosswalk

A regulatory crosswalk maps the obligations arising from Annex III(4) and Annex III(5)(a) to specific articles of the EU Charter of Fundamental Rights (European Parliament, Council of the European Union, and European Commission, 2000). This mapping is not enumerated in the AI Act itself; the crosswalk makes it explicit and machine-readable, reducing interpretive uncertainty for deployers and providing the structured link between regulatory requirements and rights evidence that the schema’s multi-label rights axis builds on.

2.5. Query Interface

A Flask web application loads the knowledge graph on startup and allows users to filter records by Annex III domain, Charter rights, and risk pattern through a form-based interface, moving the framework toward the interactive compliance tooling anticipated by Art. 27(5) (European Parliament and Council of the European Union, 2024).

3. Evaluation

3.1. Gold Standard and Agreement Metrics

Of the 150 corpus records, 69 (46%) were manually annotated to form a gold standard, stratified across sources and both Annex III domains. Each record was annotated by reviewing original source material rather than summary fields alone, following written guidelines with explicit decision criteria per axis. Cohen’s κ\kappa (Cohen, 1960) and percentage agreement are reported for each axis and method, interpreted using the Landis and Koch scale (Landis and Koch, 1977).

The 69 manually annotated records (46%) are a deliberately stratified gold standard rather than a partial annotation effort. The framework’s purpose is to test whether automated annotation can scale beyond what manual labelling feasibly covers, so the manual subset is sized to give reliable per-axis agreement estimates across both Annex III domains and all three sources, not to annotate the corpus exhaustively. Manually labelling all 150 records would defeat the evaluation, since the pipeline exists precisely to avoid that cost at scale.

3.2. Domain Classification

For the essential services domain, the best-performing method achieves κ=0.525\kappa=0.525 (moderate agreement). For employment, the best-performing method achieves κ=0.045\kappa=0.045—near-chance. This divergence is attributable to spurious vocabulary correlations: a disproportionate share of records discuss employment incidentally rather than as the primary domain of harm, and both keyword and LLM classifiers over-predict employment for records containing general labour market language.

The implication for algorithmic fairness research is direct. Employment discrimination is among the most extensively documented harms of automated decision-making, and Annex III(4) covers the AI Act’s highest-risk employment applications. Near-chance agreement means that automated annotation of employment-domain risk evidence cannot currently be trusted without human review. Deployers relying on such tools to surface employment risk evidence for a FRIA risk generating an evidence base that is both incomplete and misleading. The finding holds across models: a comparison with GPT-4o-mini yields κ=0.196\kappa=0.196 for domain classification, confirming the difficulty lies in the task rather than in a single model.

3.3. Coverage and Risk Patterns

Five FRIA demonstration scenarios query the populated knowledge graph for records relevant to realistic deployer contexts: welfare eligibility AI, public sector recruitment screening, surveillance in public housing, LLM decision-support for caseworkers, and a cross-domain thematic review by a national regulator. Across all five, 103 of the 150 records are surfaced (68.7% coverage). The 31.3% not surfaced are predominantly records with unknown classifications produced by pipeline underclassification—the direct consequence of the domain agreement failures described above. Figure 1 reports per-scenario retrieval: the welfare-eligibility and cross-domain profiling scenarios surface the largest evidence pools (57 and 50 records), while the employment-focused recruitment scenario surfaces the fewest (17), consistent with the near-chance employment agreement reported above.

Horizontal bar chart of the number of corpus records surfaced by
each of the five FRIA demonstration scenarios.
Figure 1. Records surfaced by each of the five FRIA demonstration scenarios. Scenarios may retrieve overlapping records; their union is 103 of the 150 corpus records (68.7% coverage). The employment recruitment scenario surfaces the fewest records, mirroring the near-chance employment domain classification.Horizontal bar chart of the number of corpus records surfaced by each of the five FRIA demonstration scenarios.

The hybrid method reduces unknown risk pattern classifications from 61.3% (keyword baseline) to 14.0%, substantially enlarging the evidence base available for compliance queries. Agreement on risk patterns remains in the slight-to-fair range, with procedural unfairness and lack of transparency the categories of greatest disagreement, driven primarily by multi-label harm structures where incidents involve co-occurring failure modes.

3.4. Use Case: A Welfare Eligibility Deployer

A public body preparing to deploy an AI system that supports social-welfare eligibility decisions (Annex III(5)(a)) must document, before deployment, how the system may bear on the rights to social security (Art. 34) and non-discrimination (Art. 21). Using the query interface, the deployer filters the knowledge graph for essential-services records implicating these rights; the query returns 57 candidate records (Scenario A in Figure 1). Each record carries its risk-pattern and causal-factor classifications together with full provenance back to the originating incident, so a flagged non-discrimination risk can be traced to specific precedents—such as the Dutch childcare-benefits scandal (Amnesty International, 2021) and the SyRI profiling case (District Court of The Hague, 2020; Wieringa and others, 2023)—rather than to a general assertion that such risks exist. This is what separates an evidence-grounded FRIA from a narrative one: the deployer can judge whether documented failures are analogous to its own context, and the machine-readable output lets a supervisory authority later verify that comparable deployers considered the same rights, as anticipated by Art. 27(5).

4. Limitations

Four limitations bear on the framework’s current scope. The gold standard was produced by a single annotator; multi-annotator reliability using Fleiss’ κ\kappa (Fleiss, 1971) is a necessary next step before pipeline outputs can be used without human review. The corpus covers only two of the eight Annex III categories; the schema is designed for extensibility but coverage of biometrics, law enforcement, and migration is not yet tested, a gap that prior work on high-risk AI classification highlights (Golpayegani et al., 2024a; Golpayegani et al., 2023). AIAAIC is predominantly English-language and draws heavily on U.S. and U.K. sources, introducing geographic bias in an EU regulatory context. Finally, employment classification quality is insufficient for unsupervised deployment and requires human oversight for any FRIA relying on that domain’s evidence.

5. Relevance to ECAF

This work contributes across three of ECAF’s disciplinary areas.

Policy and Law. The regulatory crosswalk and FRIA demonstration scenarios directly address impact assessment obligations under the AI Act, providing concrete, reusable infrastructure for Art. 27 compliance and contributing to emerging standardisation of AI incident reporting (OECD, 2025). This work can also be considered as a foundation for the automated tool the AI Office should develop for FRIA, as per Art. 27(5). The analysis of GDPR versus AI Act impact assessment requirements (Rintamäki et al., 2026) provides complementary legal framing.

Computer Science. The annotation pipeline, multi-method comparison, and gold standard evaluation contribute empirical evidence on the reliability of LLM-assisted classification in a regulatory context, with direct implications for auditing framework design.

Social Sciences. The corpus documents real-world AI failures in employment and essential services, two domains where automated systems bear most directly on the rights of marginalised groups. The near-chance employment domain agreement (κ=0.045\kappa=0.045) is a cautionary result for practitioners building automated fairness assessment tools in this domain.

6. Conclusion

We have presented a Semantic Web-based framework that consolidates fragmented AI risk evidence into a structured, SPARQL-queryable knowledge graph supporting EU AI Act FRIA compliance. Five FRIA demonstration scenarios achieve 68.7% coverage of a 150-record corpus, and gold standard evaluation reveals that LLM-assisted employment domain classification is near-chance—a result with direct implications for automated fairness assessment in the domain most associated with algorithmic discrimination. All schema artefacts and annotated data are released under CC BY 4.0 and pipeline source code under the MIT Licence at https://github.com/faitholopade/Dissertation, to support adoption by regulators, national authorities, and SMEs extending the framework to further Annex III categories.

Acknowledgements.
This work was supported by the School of Computer Science and Statistics, Trinity College Dublin. The authors acknowledge the support of the ADAPT Centre for Digital Content Technology, funded under the Science Foundation Ireland Research Centres Programme (#13/RC/2106_P2).

References

  • Amnesty International (2021) Amnesty International Xenophobic machines: discrimination through unregulated use of algorithms in the dutch childcare benefits scandal. Technical report Amnesty International. External Links: Link Cited by: §1, §3.4.
  • Anthropic (2024) Anthropic The Claude 3 model family: Opus, Sonnet, Haiku — model card. Note: Technical report External Links: Link Cited by: §2.3.
  • Bender et al. (2021) E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell On the dangers of stochastic parrots: can language models be too big?. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’21), External Links: Link Cited by: §1.
  • Bommasani et al. (2021) R. Bommasani et al. On the opportunities and risks of foundation models. arXiv preprint. External Links: Link Cited by: §1.
  • Cohen (1960) J. Cohen A coefficient of agreement for nominal scales. Educational and Psychological Measurement 20 (1), pp. 37–46. External Links: Document, Link Cited by: §3.1.
  • Council of Europe (2024) Council of Europe HUDOC: european court of human rights case-law database. Note: Online database External Links: Link Cited by: §2.1.
  • District Court of The Hague (2020) District Court of The Hague Judgment in the case of njcm c.s. versus the state of the netherlands (syri case), case no. c-09-550982. Note: Judgment of the District Court of The Hague, 5 February 2020 External Links: Link Cited by: §1, §3.4.
  • European Parliament and Council of the European Union (2024) European Parliament and Council of the European Union Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Note: Official Journal of the European Union External Links: Link Cited by: §1, §2.5.
  • European Parliament, Council of the European Union, and European Commission (2000) European Parliament, Council of the European Union, and European Commission Charter of fundamental rights of the European Union. Note: Official Journal of the European Communities, C 364/1, 18.12.2000 External Links: Link Cited by: §1, §2.4.
  • European Union Agency for Fundamental Rights (2025) European Union Agency for Fundamental Rights Assessing high-risk artificial intelligence: fundamental rights risks. Technical report FRA. External Links: Link Cited by: §1.
  • Fleiss (1971) J. L. Fleiss Measuring nominal scale agreement among many raters. Psychological Bulletin 76 (5), pp. 378–382. External Links: Document, Link Cited by: §4.
  • Golpayegani et al. (2022) D. Golpayegani, H. J. Pandit, and D. Lewis AIRO: AI risk ontology (v1). Note: Working paper External Links: Link Cited by: §1, §2.2.
  • Golpayegani et al. (2023) D. Golpayegani, H. J. Pandit, and D. Lewis To be high-risk, or not to be—semantic specifications and implications of the AI act’s high-risk AI applications and harmonised standards. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’23), pp. 905–915. External Links: Document, Link Cited by: §4.
  • Golpayegani et al. (2024a) D. Golpayegani, H. J. Pandit, and D. Lewis To be high-risk, or not to be—semantic specifications and implications of the AI act’s high-risk AI applications and harmonised standards. Note: Working paper / technical report External Links: Link Cited by: §4.
  • Golpayegani et al. (2024b) D. Golpayegani, H. J. Pandit, and D. Lewis VAIR: vocabulary of AI risks. External Links: Link Cited by: §1, §2.2.
  • Landis and Koch (1977) J. R. Landis and G. G. Koch The measurement of observer agreement for categorical data. Biometrics 33 (1), pp. 159–174. External Links: Document, Link Cited by: §3.1.
  • McGregor (2021) S. McGregor Preventing repeated real world AI failures by cataloging incidents: the AI incident database. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35, pp. 15458–15463. External Links: Link Cited by: §1.
  • OECD (2025) OECD Towards a common reporting framework for AI incidents. Technical report OECD Artificial Intelligence Papers, No. 34. External Links: Link Cited by: §5.
  • Office of Management and Budget (2024) Office of Management and Budget 2024 consolidated federal AI use case inventory. Note: Published pursuant to OMB Memorandum M-24-10 and the AI in Government Act of 2020 External Links: Link Cited by: §2.1.
  • Olopade (2026) F. Olopade A reusable Semantic Web-based framework linking LLM risks to fundamental rights for EU AI Act high-risk public sector applications. Master’s Thesis, Trinity College Dublin. External Links: Link Cited by: §1.
  • Pandit et al. (2024) H. J. Pandit, B. Esteves, G. P. Krog, P. Ryan, D. Golpayegani, and J. Flake Data privacy vocabulary (DPV)—version 2. Note: arXiv preprint arXiv:2404.13426 / ISWC 2024 External Links: Link Cited by: §1, §2.2.
  • Pownall (2024) C. Pownall AIAAIC: ai, algorithms, and automation incidents and controversies. External Links: Link Cited by: §1, §2.1.
  • Rintamäki et al. (2026) T. Rintamäki, D. Golpayegani, D. Lewis, E. Celeste, and H. J. Pandit Impact assessment requirements in the GDPR vs the AI act: overlaps, divergence, and implications. Note: OSF Preprints External Links: Document, Link Cited by: §5.
  • Rintamäki and Pandit (2025) T. Rintamäki and H. J. Pandit Developing an ontology for AI act fundamental rights impact assessments. Note: arXiv preprint arXiv:2501.10391 External Links: Link Cited by: §2.2.
  • W3C Data Privacy Vocabulary Community Group (2022) W3C Data Privacy Vocabulary Community Group Data privacy vocabulary (dpv) specification. External Links: Link Cited by: §1, §2.2.
  • Weidinger et al. (2021) L. Weidinger et al. Ethical and social risks of harm from language models. arXiv preprint. External Links: Link Cited by: §1.
  • Wieringa et al. (2023) M. Wieringa et al. “Hey syri, tell me about algorithmic accountability”: lessons from a landmark case. Data & Policy. External Links: Document, Link Cited by: §1, §3.4.