Nesi1 — Uncensored AI for Cybersecurity

Nesi1 is an uncensored large language model, curated and optimized by Nesilabs specifically for offensive and defensive cybersecurity: penetration testing, vulnerability analysis, exploit development, secure code review, and audit automation.

It is built on the Qwen3.8-27B architecture and fine-tuned by Nesilabs on a curated security corpus for real-world workflows and autonomous security agents.

Highlights

  • Unrestricted for legitimate security use cases.
  • 27B parameters, long context (up to 256k tokens), native tool-calling.
  • Optimized for pentest agents and security copilots.
  • Fine-tuned on a curated security corpus (see Training).

Training

Nesi1 is specialized in offensive/defensive security, trained on a curated corpus of ~9.7k full-length technical documents:

  • Disclosed HackerOne reports
  • Bug bounty writeups
  • HackTricks
  • MITRE ATT&CK
  • Nuclei templates
  • Exploit-DB exploits

Plus 5,050 source-anchored Q/A pairs — each verified against its source document, with 0 benchmark contamination.

Model details

Base architecture Qwen3.8-27B (hybrid Qwen3-Next)
Quantization AWQ 4-bit (W4A16), ~18 GB
Context length up to 262,144 tokens
Modality text + vision (multimodal)
Tool-calling OpenAI-compatible (qwen3_coder parser)
Serving SGLang / vLLM compatible

Deployment

Nesi1 runs on any OpenAI-compatible inference server (SGLang, vLLM). This is the exact SGLang configuration we run in production:

python3 -m sglang.launch_server \
  --model-path nesilabs/Nesi1 \
  --served-model-name Nesi1 \
  --context-length 262144 \
  --mem-fraction-static 0.92 \
  --trust-remote-code \
  --kv-cache-dtype auto \
  --tool-call-parser qwen3_coder \
  --default-chat-template-kwargs '{"enable_thinking":false}' \
  --host 0.0.0.0 --port 30000

Then query it like any OpenAI chat endpoint (POST /v1/chat/completions).

What we learned (tips that matter)

  • --tool-call-parser qwen3_coder is required for clean OpenAI-style tool_calls. Without it, tool calls come back as raw <tool_call> text and break agent frameworks.
  • enable_thinking: false keeps internal reasoning out of the final answer.
  • bf16 KV cache (--kv-cache-dtype auto) for quality. FP8 KV roughly doubles capacity but we observed quality drift on long agentic sessions, so we keep bf16.
  • Prefix caching (RadixCache) is the big win for agents. In agentic loops the same context is replayed every turn — cache-read hit rates of ~95%+ make prefill almost free. Keep the agent context append-only to maximize hits.
  • Speculative decoding (NEXTN) did not help on this model under concurrency — it lowered throughput, so we left it off.

Performance (measured on 1× RTX PRO 6000 Blackwell 96GB, bf16)

Metric Value
Single-stream decode ~84 tok/s
Under concurrency (several streams) ~50–75 tok/s per stream
KV cache pool ~595k tokens
Concurrent sessions @ ~60k ctx ~10
Concurrent multi-agent pentest pipelines ~2–3
Max context up to 256k tokens

A single 96GB GPU comfortably serves a team's security copilots plus a few parallel pentest pipelines. For higher concurrency, load-balance across GPUs.

⚠️ Legal Notice / Disclaimer

Nesi1 is intended EXCLUSIVELY for security professionals, for lawful and authorized activities.

  • Authorization required. Use it only against systems, networks, or applications for which you have explicit written permission (a contracted penetration test, an in-scope bug bounty program, or your own/lab environments). Using it against systems without authorization is illegal and strictly prohibited.
  • User responsibility. You are the sole responsible party for how you use this model and for complying with all applicable laws in your jurisdiction (including computer-misuse, data-protection/GDPR, and intellectual-property laws).
  • No warranty. The model is provided "as is", without warranty of accuracy, fitness, or results. It may produce incorrect information — always verify before acting.
  • Limitation of liability. Nesilabs is not liable for any damages, losses, or legal consequences arising from the use or misuse of this model, including any unauthorized or unlawful use.
  • Prohibited use. Using Nesi1 for illegal activities, unauthorized access, extortion, harm to third parties, or any purpose that violates the law or these terms is prohibited.

By using Nesi1 you accept this notice and the Nesilabs Terms of Service and Acceptable Use Policy.

License

Released under Apache 2.0. Based on the Qwen3.8-27B architecture. Adapted and packaged by Nesilabs.

Downloads last month
29
Safetensors
Model size
27B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support