Building multimodal benchmarks to evaluate the accessibility of GCSE and A-level assessment design
Yulan He, King’s College London
Summary: This project will develop an open multimodal benchmark for evaluating whether AI systems can identify and help address accessibility barriers in GCSE and A-level assessment materials. Working with AQA, the team will curate and annotate examination items that combine text, layout, colour, diagrams and images. The benchmark will include tasks for detecting accessibility problems and proposing revisions that improve accessibility without changing what a question is intended to assess. It will release benchmark data and annotations where licensing permits, together with reproducible baselines, an evaluation harness, documentation and community-facing evaluation resources. The project aims to support fairer assessment design for students with special educational needs, English as an additional language, and other diverse backgrounds.
MAICA-Bench: A Multimodal AI Compliance and Adversarial Auditing Benchmark for Regulated-Sector Deployment
Mohamed Chahine Ghanem, Keele University
Summary: MAICA-Bench will create an open benchmark for testing the robustness, compliance evidence and deployment readiness of multimodal AI systems used in regulated sectors. In collaboration with PRIAM Cyber AI, it will develop a shared adversarial evaluation harness spanning security telemetry, analyst notes and text, audio, screenshots and other visual evidence, including tests of contradictions across modalities. The benchmark will connect technical results to assurance and regulatory requirements without claiming to provide certification. Planned outputs include reusable tasks and scenarios, evaluation code, reproducibility manifests, documentation and a public leaderboard, with data released as openly as privacy and licensing allow. The benchmark is intended for researchers, security operations teams, auditors and AI-governance practitioners in areas such as finance, cybersecurity and operational technology.
Deployment-Centric Multimodal AI Benchmark for Transparent Organic Photovoltaic Materials Development
Alex Ramadan, University of Sheffield
Summary: This project will build an open, deployment-centred benchmark for AI-assisted development of transparent organic photovoltaic materials, including applications in agrivoltaics. It will combine complementary information about molecular and chemical structure, material and device properties, processing metadata, optical and electrical measurements, and practical factors such as cost and supply-chain risk. Building on existing open organic photovoltaic resources and the team’s RealMat-BaG experience, and supported by experimental data and expertise from TerraChange Solar, the benchmark will test models under realistic conditions such as out-of-distribution splits, incomplete modalities, uncertainty and interpretability requirements. Outputs will include curated data, reproducible baselines, evaluation tools, documentation and leaderboard support for the wider AI-for-materials community.
CollabSim: A Multimodal Benchmark for Scalable Human-Robot Collaboration
Oya Celiktutan, King’s College London
Summary: CollabSim will develop a scalable multimodal benchmark for evaluating human-robot collaboration in realistic and safety-relevant settings. The benchmark will bring together motion, physical interaction and force information, language and task context, using both interaction data and controllable simulation. Generative human-behaviour modelling and physics-based constraints will be used to expand the range of scenarios, including difficult and failure cases that are costly or unsafe to collect repeatedly in the real world. Working with Honda R&D Europe and Guy’s and St Thomas’ NHS Foundation Trust, the team will define and validate meaningful collaborative tasks. Planned outputs include open benchmark assets, evaluation protocols, baseline models, simulator-integration tools and documentation to support reproducible comparison of collaboration, robustness, safety and generalisation.
VeracityBench: An Open Benchmark for Evaluating Multimodal AI Video Verification
Mohsen Mosleh, University of Oxford
Summary: VeracityBench will create an open benchmark for evaluating how multimodal AI systems assess the credibility of online video and communicate the evidence and uncertainty behind their judgements. The project will curate a diverse set of real-world videos and contextual information, with expert verification, structured annotations and public-perception data. AI outputs will be compared with professional fact-checker assessments and crowd responses, allowing the benchmark to measure not only whether a conclusion is correct, but also whether the system uses appropriate evidence and is well calibrated when uncertain. The benchmark will be co-designed with BBC Verify, Full Fact, Indicator and the United Nations Information Integrity Unit. Code, task definitions, annotation and evaluation tools, documentation, and versioned data or metadata will be released as openly as copyright and platform rules permit.
OMAIB-Dementia: An Open, Federated Multimodal Benchmark for Early Dementia Diagnosis
Zahara Gironés Delgado-Ureña, University of Cambridge
Summary: OMAIB-Dementia will create an open and federated benchmark for evaluating multimodal AI systems for early dementia diagnosis. It will combine MRI sequences, cognitive assessments, clinical records and radiology text, with an open track based on public datasets and a secure federated track using NHS memory-clinic data that remain at their host institutions. The benchmark will evaluate clinically meaningful performance, including discrimination at relevant operating points, calibration, robustness to missing modalities, fairness and privacy-preserving execution. A distinctive aim is to compare central and federated evaluation on the same clinical cases, providing evidence about whether distributed evaluation faithfully reflects real-world performance. Working with clinicians and Flower Labs, the project will release tasks, harmonisation and evaluation software, reference baselines, documentation and a public leaderboard.
HabiBench: An Open Multimodal Benchmark for Reliable UK Terrestrial Habitat Assessment
Hongrui Shi, University of Lincoln
Summary: HabiBench will develop a fully open multimodal benchmark for reliable terrestrial habitat assessment in the UK. It will harmonise ground-level habitat photographs, Earth-observation embeddings and ecological text or prior information using open resources including LUCAS and TESSERA. Rather than reporting classification accuracy alone, the benchmark will test which modalities contribute to each decision, how well uncertainty is calibrated, how performance changes when a modality is missing, and when an AI system should support, triage or defer to a field survey. The benchmark will be validated with Mozaic Earth and informed by ecological end users, including UKCEH expertise. Outputs will include the dataset, a reusable construction pipeline, baselines, datasheets and model cards, a reliability report, documentation and leaderboard integration, with potential extension to wider European habitat monitoring.
Contact us: ukomain.contact@gmail.com