Strong AI signal
99.8%of 2,500 synthetic images warnedAccuracy
What it catches.
What it misses.
Every number below was measured on data the deployed model never trained on, and is reported with its sample size, its false-positive rate and the cases where it fails. Measured 4 August 2026 on aidetect-v5-image-384-v1.
Images: a fresh set of 8,133.
Assembled on the day of the measurement so that no part of it could have leaked into training. 5,633 real photographs from five sources, 2,500 synthetic images from four generator families the model had never seen.
False alarms
1.90%of 5,633 real photographs in this benchmark, at most 2.23% with 95% confidence — but see the correction belowPrecision
95.9%of strong warnings were in fact synthetic| Tier | Synthetic recall | Real false positives | 95% upper bound | Precision |
|---|---|---|---|---|
| Strong AI signal | 99.8% (2,495/2,500) | 1.90% (107/5,633) | 2.23% | 95.9% |
| Possible AI signal | 99.9% (2,498/2,500) | 4.12% (232/5,633) | 4.58% | 91.5% |
| Source | False positives | Images |
|---|---|---|
| Smithsonian portraits | 0.00% | 302 |
| Smithsonian art and scans | 0.11% | 890 |
| Camera-native and stock | 1.38% | 3,198 |
| NASA video posters | 3.06% | 294 |
| NASA imagery | 5.58% | 949 |
| Generator | Recall | Images |
|---|---|---|
| Playground 2.5 | 100.0% | 625 |
| SANA 1.6B | 100.0% | 625 |
| Shuttle 3 | 100.0% | 625 |
| AuraFlow 0.3 | 99.2% | 625 |
Correction, 4 August 2026: it false-alarms on edited photography.
The 1.90% above is real, but it is measured on the five sources in that benchmark: camera-native stock, NASA imagery and Smithsonian scans. Those are plain photographs. They do not represent heavily post-processed photography, and we found out the hard way that the detector treats that as synthetic.
On a sample of 37 Wikimedia Commons featured pictures — peer-reviewed, human-made photographs and artwork reproductions — the detector raised a strong false alarm on 24 of them, 64.9% (95% confidence interval 47.5% to 79.8%). The same detector false-alarms on 1.38% of camera-native stock photographs in the benchmark. That is not a small difference in degree; it is a different failure regime.
The likely cause is a gap on the real side of training rather than the synthetic side. Everything the model learned to call synthetic is latent text-to-image output, which is sharp, well-composed, richly lit and aesthetically polished. Featured pictures are, by definition, the most polished real photographs available. The model appears to have learned that polish means synthetic.
What this means for you. If you check an ordinary snapshot, a screenshot or a stock photograph, the numbers above apply. If you check an award-winning landscape, a professional edit or a fine-art reproduction, expect a false alarm and treat a strong signal on such an image as close to meaningless until this is fixed. We are publishing this rather than waiting for a fix, because the page was live with the better number on it.
Where it is blind.
All four generators above are latent text-to-image models — the same family the detector was trained on. Near-perfect recall there says nothing about other kinds of generator, and on a benchmark that contains them the picture changes completely.
| How the image was generated | Recall |
|---|---|
| Latent text-to-image (Midjourney, SD 1.5, Wukong) | 50.4% |
| Pixel diffusion (GLIDE) | 9.8% |
| GAN, class-conditional (BigGAN) | 7.0% |
| VQ diffusion (VQDM) | 3.6% |
| Pixel diffusion, class-conditional (ADM) | 2.5% |
The axis is not how recent a generator is — it is how it works. Feeding the model 3,200 more images from 2023 moved nothing, because those were latent text-to-image too. Closing this gap needs GAN, pixel-diffusion and VQ examples in training, and that work is not done.
Video: one generator solved, three not.
740 frames across six sources, aggregated to one verdict per clip. The lane samples frames and scores them as images — it does not model motion, temporal consistency or lip-sync, and it is not a deepfake detector.
| Source | Clips warned | Previous detector | Clips |
|---|---|---|---|
| Kling 2.1 | 100.0% | 29.3% | 300 |
| Text2Video-Zero | 33.3% | 91.7% | 60 |
| Pika | 20.0% | 83.3% | 60 |
| Sora | 13.3% | 66.7% | 60 |
| Real video (false alarms) | 0.0% | 5.0% | 60 |
| NASA video posters (false alarms) | 5.5% | 6.5% | 200 |
The current detector turned the previous one’s worst video case into its best, and its best cases into its worst. Both halves of that sentence are on this page on purpose.
What we deliberately do not claim.
- No comparison to other detectors. We have not benchmarked competitors on these sets, so there is no “most accurate” claim anywhere on this site.
- No single accuracy percentage. A recall figure without its false-positive rate, its threshold and its generator mix is not a fact about a detector.
- Never “authentic”. A low score is the absence of a signal, not evidence that a human made something. The result vocabulary has no value that means authentic.
- It false-alarms on polished photography. 64.9% of a 37-image sample of Wikimedia featured pictures drew a strong warning. The benchmark number does not cover that kind of image, and we say so above rather than only in a footnote.
- It misses our own bar. Our internal release gate is a false-positive rate of 0.6% or below; at 1.90% this lane is above it. It ships because a warning carries no automatic consequence — and that is exactly why no headline percentage appears at the top of this page.
- Text and audio are not in the product. They measured 79% accuracy at 13% false positives, and 84% balanced accuracy on 193 clips. Too weak to publish.
How to read a result.
A strong warning on an image from a mainstream generator is reliable — that is the case these numbers cover. A quiet result is much weaker evidence: it can mean the image is real, or that it came from a kind of generator this detector cannot see.
Content Credentials outrank the model. When a file carries a verified C2PA assertion or metadata naming an AI tool, that is evidence about the file itself rather than a statistical guess, and the checker shows it separately.