FARE: The Forensic Watchdog That Catches AI Providers Swapping Certified Image Generators
FARE detects silent AI generator swaps at deployment time using only output images — no model access required. NeurIPS 2026 research from Edinburgh.

Main Story
When an enterprise or regulator certifies an AI image generator for deployment in a high-stakes domain, they typically inspect the model once — then hand control back to the provider. What happens next is largely a matter of trust. A new forensic auditing framework called FARE (Forensic Acceptance Region Estimation), developed by researchers Kai Yao and Marc Juarez at the University of Edinburgh's School of Informatics, is designed to remove that trust assumption entirely.
The problem FARE addresses is structurally straightforward but technically difficult to solve. As the research paper notes, modern AI image generators are increasingly deployed as opaque APIs, where customers can query the deployed service but cannot inspect model weights or architecture. This creates a meaningful integrity gap: a provider may pass governance certification with one generator and later silently switch to a cheaper and lower-quality one for deployment, compromising public trust or even safety in high-stakes domains — including, as the paper notes, defence and healthcare contexts where generative AI is increasingly being considered.
Regulatory pressure is sharpening this concern. The EU AI Act establishes conformity-assessment requirements for high-risk AI systems, and the paper frames FARE directly as a technical response to a question those requirements raise but do not mechanically answer: how can an external auditor or customer verify that an output from a deployed service truly came from the certified generator?
FARE's approach formalises the problem in two distinct phases: enrollment and verification. During enrollment, an auditor certifies a generator under an NDA-like legal agreement and enrolls it into FARE by sampling images. This phase follows governance certification, assumes provider collaboration, and is codified in a contract specifying the model checkpoint and inference pipeline settings. Models deviating from the contract — for instance, lower-quality models — are treated as non-certified.
During verification, the auditor receives only a queried image from the deployed generator and determines whether it is consistent with the certified generator. Crucially, this check requires only that single image. No model access, no parameter inspection, no side-channel queries — only pixels.
The work is accepted for publication in the proceedings of the 40th Annual Conference on Neural Information Processing Systems (NeurIPS 2026), and was supported by the Edinburgh International Data Facility (EIDF) and the Data-Driven Innovation Programme at the University of Edinburgh. Co-author Marc Juarez is a GAIL Fellow and recipient of a Google Research Scholar Award in Security.
FARE builds on a prior line of research by the same team. Their earlier study, Smudged Fingerprints (presented at IEEE SaTML 2026), systematically evaluated the robustness of existing AI image fingerprinting methods and found that fingerprint removal attacks often achieved success rates above 80% in white-box settings, with none of the evaluated techniques delivering both high accuracy and resistance to attacks across all threat scenarios. FARE was designed with those attack findings explicitly in mind.
Technical Breakdown
Forensic Feature Basis: FARE's features are based on image generator-specific artifacts that have been proposed for forensic applications — subtle structural traces that different model architectures and training configurations leave in their outputs.
Enrollment Architecture: A constrained convolutional front end and patch scoring mechanism capture generator traces. During enrollment, FARE trains using only images from the single certified generator; no images from alternative or substitute generators are required at this stage.
Hard-Sample Amplification: FARE amplifies its forensic features during training by finding hard samples — adversarial and contradiction positives — that tighten the acceptance region and increase sensitivity to subtle changes between certified and non-certified generators.
Swap Detection Scenarios: The evaluation covers four swap scenarios: cross-family swaps using an FFHQ-256 generator pool and a diffusion-dominated CommunityForensics pool, and within-family swaps covering StyleGAN training configurations and Stable Diffusion versions. The within-family scenarios are the more demanding, as they require detecting substitutions between closely related model variants.
Detection Performance: Across all four swap scenarios, FARE achieves 92.56–99.45% scenario mean true positive rate (TPR) at 1% false positive rate (FPR) — a strict operating point designed to minimise false alarms in high-stakes deployment contexts.
Adversarial Robustness: Under a perceptually constrained white-box projected gradient descent (PGD) attack — with an L∞ perturbation budget of 0.025 and LPIPS constraint below 0.05 — scenario mean TPR remains 92.38–99.38%, demonstrating that the acceptance region resists the evaluated exact-model and decision-only adversarial conditions.
Deployment Constraint: FARE operates with zero access to the deployed model's weights or architecture, making it compatible with the opaque-API deployment model that dominates the commercial market.
Industry Impact
For AI Governance and Certification Bodies: FARE offers the first published, formally evaluated mechanism for post-certification, deployment-time integrity auditing of image generators. Standards bodies developing conformity-assessment frameworks under regulations such as the EU AI Act now have a concrete technical approach to reference for the verification phase of an audit lifecycle — a phase that existing frameworks largely leave to provider self-reporting.
For Enterprise API Consumers: Any organisation procuring image generation capabilities via an opaque API — from media production to medical imaging to satellite image synthesis — faces the principal-agent problem FARE is designed to solve. The method requires only query access, meaning a buyer can run ongoing integrity checks without renegotiating API terms or gaining model access.
For AI Service Providers: FARE's existence changes the calculus around undisclosed model substitution. Providers operating in regulated or high-stakes sectors will need to ensure that deployment pipelines maintain the exact certified checkpoint and inference configuration, as the method is sensitive to within-family swaps including version differences within architectures such as Stable Diffusion.
For Forensic and Security Tool Developers: The open research release via arXiv and GitHub (kaikaiyao/FARE) provides a reproducible baseline. However, the Smudged Fingerprints findings that motivated FARE's design also serve as a caution: the adversarial robustness of any forensic method must be treated as a moving target. FARE's hard-sample training approach is a direct engineering response to that fragility, but real-world deployment will require continued red-teaming as evasion techniques evolve.
For the Broader Generative AI Supply Chain: As generative AI is increasingly considered in safety-critical domains such as defence and healthcare, the gap between governance certification and deployment-time verification becomes a material risk. FARE establishes a technical template for closing that gap — shifting integrity auditing from a one-time compliance event into a continuous, evidence-based process.
