You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Blue-Eye is a content-moderation model. By downloading it you agree to use it under the DINOv3 License and not to use it to monitor or make decisions about individual people.

Log in or Sign Up to review the conditions and access this model content.

Blue-Eye logo

Model Card for Blue-Eye

Blue-Eye is an image classifier for content moderation. Given an image, it predicts whether the content is safe, suggestive or explicit, together with a probability for each class. On a benchmark of 3,000 real photographs it reaches 88.9% accuracy, ahead of AWS Rekognition, Gemini 3.1 Pro, Google Cloud Vision and the open-source NSFW detectors it was compared with. It handles both photographs and anime/illustration.

Overall accuracy on 3,000 real photographs

Model Details

Blue-Eye is a DINOv3 ViT-L/16 vision transformer fine-tuned end to end for three-class content classification. The model takes a 512x512 RGB image and returns three class probabilities.

Model Description

Model Sources

Uses

Direct Use

Blue-Eye is built for moderating sexual content in images:

  • filtering explicit or suggestive images out of feeds, search results and timelines
  • blurring images or adding content warnings
  • age-gating content on platforms that allow adult material
  • prioritising images for human moderators
  • curating image datasets before training other models

The class definitions follow a nudity and sexual-content rubric:

  • safe: no sexualised content, including swimwear, fitness, medical images, breastfeeding and non-sexual art
  • suggestive: sexualised but not explicit, such as posed lingerie shots or bare buttocks
  • explicit: exposed genitalia, sexual acts or full nudity

Downstream Use

The default prediction is the highest-probability class. Platforms with a stricter or looser policy can set their own threshold on p(explicit) or p(suggestive) + p(explicit), tuned on their own data. The model can also be fine-tuned further on a platform's own labels.

Out-of-Scope Use

Blue-Eye covers sexual content only; violence, gore and other policy areas are out of scope. It does not estimate age and is not a tool for detecting child sexual abuse material, for which dedicated hash-matching services should be used. It should not be used to monitor or make decisions about individual people.

Bias, Risks, and Limitations

The boundary between suggestive and its neighbouring classes is the most subjective part of the task, for people and models alike, and it is where most disagreements occur. Performance across demographic groups was not part of this evaluation.

Recommendations

Validate the model on images representative of your own platform before deployment, and keep a human review step for actions that affect user accounts.

How to Get Started with the Model

pip install torch "transformers>=4.56" safetensors pillow numpy huggingface_hub
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("pp1618/Blue-Eye")
sys.path.insert(0, path)

from inference import classify

for result in classify(["photo.jpg", "drawing.png"], model=path):
    print(result["label"], result["probabilities"])

From the command line:

python inference.py photo.jpg folder_of_images/ --model pp1618/Blue-Eye --device cuda

On GPUs with bfloat16 support, add --precision bf16 (or precision="bf16" in Python) for faster inference with practically identical predictions.

Training Details

Training Data

About 2.6 million web images covering real photographs and anime/illustration, labelled into the three classes using a commercial content-moderation service, source content ratings and model-assisted relabelling. Evaluation images were removed from the training data.

Training Procedure

Training ran in three progressive fine-tuning stages starting from the DINOv3 ViT-L/16 checkpoint:

  1. about 1.0M real photographs
  2. about 1.0M images combining anime/illustration with real photographs
  3. about 650k class-balanced images

Training regime: 2 epochs per stage at 512x512, AdamW with a one-cycle schedule, label smoothing 0.05, random resized crops, horizontal flips and colour jitter, float32 master weights with bfloat16 mixed precision.

Evaluation

Testing Data and Metrics

The Blue-Eye benchmark contains 3,000 real photographs (1,480 safe, 513 suggestive, 1,007 explicit), including 525 non-sexual, skin-heavy images such as swimwear, fitness and medical photos. Every system below was evaluated on the same images with the same three-class labels; commercial services were queried in August 2026 and their outputs mapped to the three classes. Google Cloud Vision counts an image as explicit when adult is VERY_LIKELY and as suggestive when racy is VERY_LIKELY; AWS Rekognition uses its default 50% confidence. The metric is three-class accuracy.

Results

System Type Accuracy
Blue-Eye open weights 88.9%
AWS Rekognition commercial API 87.3%
Gemini 3.1 Pro commercial model 86.4%
Gemini 3.7 Flash commercial model 84.1%
Google Cloud Vision SafeSearch commercial API 83.5%
TostAI/nsfw-image-detection-large open weights 79.2%
Marqo/nsfw-image-detection-384 open weights 73.8%
Falconsai/nsfw_image_detection open weights 70.9%
NudeNet open weights 70.4%
Freepik/nsfw_image_detector open weights 67.6%
AdamCodd/vit-base-nsfw-detector open weights 62.1%

Many open-source detectors are binary, so they were also compared on the two binary tasks:

System Safe vs not safe Explicit vs rest
Blue-Eye 92.2% 95.1%
Marqo/nsfw-image-detection-384 86.4% 78.3%
TostAI/nsfw-image-detection-large 85.7% 90.1%
Falconsai/nsfw_image_detection 83.3% 75.6%
Freepik/nsfw_image_detector 82.6% 71.8%
AdamCodd/vit-base-nsfw-detector 78.6% 62.7%
NudeNet 76.9% 87.5%

Blue-Eye results across domains:

Evaluation set Images Accuracy
Real photographs 4,000 91.4%
Anime / illustration 2,000 87.2%
All 6,000 90.0%

Per-class recall on the 3,000 real photographs: safe 89.5%, suggestive 75.4%, explicit 94.8%.

Technical Specifications

Model Architecture

  • Backbone: DINOv3 ViT-L/16: 24 layers, embedding dimension 1024, 16 heads, 4 register tokens, RoPE
  • Pooling: class token concatenated with the mean of the remaining output tokens (2048 features)
  • Head: LayerNorm followed by a linear layer to 3 classes
  • Input: RGB, shorter edge resized to 537 (bicubic), centre crop to 512x512, ImageNet normalisation
  • Weights: float32 safetensors; the head always runs in float32, including under bfloat16 inference

Compute Infrastructure

  • Hardware: Google Cloud TPU v5e
  • Software: PyTorch, Hugging Face Transformers

License

Blue-Eye is a derivative of DINOv3 and is released under the DINOv3 License. The full text is in LICENSE. Commercial use is permitted under its terms.

Citation

BibTeX

@misc{patel2026blueeye,
  title  = {Blue-Eye: a content-safety image classifier},
  author = {Patel, Pranshu},
  year   = {2026},
  url    = {https://huggingface.co/pp1618/Blue-Eye}
}
Downloads last month
3
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pp1618/Blue-Eye