Pinkerton AI

    Dolphin Mistral 24B Venice Edition: Specs, Story, and Where to Use It Now

    If you've seen dolphin-mistral-24b-venice-edition on Hugging Face or Reddit and want to actually talk to it, you have two working options in 2026: run a quantised copy on your own hardware, or use its hosted successor online. Here's the model, its story, and both paths.

    Chat anonymously
    100% anonymous·Zero logs·Encrypted

    Dolphin Mistral 24B Venice Edition is an uncensored fine-tune of Mistral's 24B model, built by the Dolphin team (Cognitive Computations) for Venice AI. The original API slug has been retired; its direct successor, Venice Uncensored 1.2, is what Venice serves today — and you can chat with it at Pinkerton AI without creating an account. For local use, GGUF quantisations of the original remain available on Hugging Face and run under Ollama.

    Last updated

    request_flow.diagram3 steps
    REQUEST
    Your prompt
    Any topic, any tone — no pre-filter.
    PIPELINE
    No moderation layer
    Skips the refusal/disclaimer pass mainstream models run.
    RESPONSE
    Direct, complete answer
    Nothing held back, nothing logged.

    What is Dolphin Mistral 24B Venice Edition?

    A collaboration between Cognitive Computations — the team behind the Dolphin series of uncensored fine-tunes — and Venice AI, applied to Mistral's 24B base. Dolphin models are trained to follow the system prompt absolutely and to answer without refusal boilerplate, which made the Venice Edition one of the most-used uncensored models of its generation: strong at creative writing, roleplay and blunt technical answers, small enough to self-host, permissively licensed.

    Why can't you find it on the Venice API anymore?

    Venice retires model slugs as newer versions ship, and the dolphin-mistral-24b-venice-edition slug has been removed from api.venice.ai's model list. Its lineage continues as Venice Uncensored — currently version 1.2 — which serves the same role as the platform's flagship uncensored fine-tune, at fp16 with a 128K context window. If a tutorial or Reddit thread points you at the old slug, Venice Uncensored 1.2 is what it resolves to in practice.

    How to use it online without an account

    Pinkerton AI hosts Venice Uncensored 1.2 in its model picker — open the chat, select it, and talk. No email, no sign-up, and no server-side conversation logs; a free session includes enough credits to evaluate it properly. This is the fastest path from 'saw the model on Hugging Face' to actually using it, and the privacy posture — anonymous sessions, crypto payment if you upgrade — matches what draws people to Dolphin models in the first place.

    How to run the original locally

    GGUF quantisations of dolphin-mistral-24b-venice-edition are published on Hugging Face and load in Ollama, LM Studio or llama.cpp. Budget roughly 14 GB for a Q4 quantisation and materially more for Q8; a 24 GB GPU gives comfortable speed, and CPU-only inference works but is slow. Local wins on absolute privacy — the prompt never leaves your machine — at the cost of quantisation quality loss and hardware. It's the right call when your threat model demands it.

    Is it actually uncensored?

    Yes, in the specific Dolphin sense: the fine-tune removes refusal behaviour and defers to your system prompt, so legal-but-sensitive requests — dark fiction, explicit roleplay between adults, frank security detail, unhedged opinions — get direct answers. It does not make illegal content available, and hosted access on Pinkerton AI enforces the same legal line regardless of model.

    Pinkerton AI vs Running it locally: how do they compare?

    FeaturePinkerton AIRunning it locally
    Model versionVenice Uncensored 1.2 (successor)Original 24B GGUF from Hugging Face
    SetupNone — select it and chatOllama / llama.cpp + ~14-50 GB download
    Precisionfp16, 128K contextQuantised (Q4-Q8), shorter practical context
    Hardware neededNone24 GB+ VRAM for comfortable speed
    AccountNone — anonymous sessionNone
    PrivacyZero server-side logsFully local — strongest possible

    FAQ

    What replaced Dolphin Mistral 24B Venice Edition?

    Venice Uncensored 1.2 — the continuation of the same uncensored flagship line on Venice's infrastructure. It's selectable at Pinkerton AI with no account.

    Can I still download Dolphin Mistral 24B Venice Edition?

    Yes — GGUF quantisations remain on Hugging Face (dphn/dolphin-mistral-24b-venice-edition and community GGUF repos) and run under Ollama, LM Studio or llama.cpp.

    What hardware does Dolphin Mistral 24B need locally?

    Around 14 GB of memory for a Q4 GGUF, more for higher quantisations; a 24 GB VRAM GPU runs it comfortably. Hosted access needs nothing.

    Is Dolphin Mistral good for roleplay?

    It's one of the reference models for it — the Dolphin training makes it follow persona and scenario instructions from the system prompt without breaking character to refuse.

    Can I try the Dolphin/Venice lineage free online?

    Yes — Pinkerton AI's free tier includes credits and requires no account; Venice Uncensored 1.2 costs 3 credits per message.

    Start chatting — free to try

    No signup. No email. Just open the chat.