Deploying an Uncensored Cloud LLM

It has long been accepted in the infosec community that open source exploits do more to help defenders than to hurt. If only malicious actors have exploits and tooling then vulnerabilities cannot be detected with precision and actionable results. The world of AI/LLM has generally not caught up to this. Legitimate penetration testers are often fighting with LLMs that flag their security testing prompts as malicious and hamper operations.
I myself have grown tired of trying to use clever prompts to convince LLMs to assist in Red Team exploits. For that reason, I finally found time to spin up my own open source uncensored/abliterated LLM in the cloud that is generally affordable and reasonably easy to setup. This LLM answers easily to prompts regarding technical assistance with penetration testing. Below are my specs and steps for setup.
- Hosting: runpod.io 'secure cloud' container
- Cost: $0.54/hr total
- GPU: RTX A6000 (48 GB VRAM)
- vCPU: 16 vCPUs on an AMD EPYC 7543
- RAM: 62 GB
- Container disk: 20 GB
- Volume disk: 40 GB, mounted at /workspace (18 GB used after initial build)
- GPU driver: 595.91.07
- CUDA version: 13.2
- Container image: ghcr.io/ggml-org/llama.cpp:server-cuda
- Template: Qwen3.8-27B-Uncensored-Q4_K_M (24GB VRAM, 192k context, OpenAI API, llama.cpp) — Template ID 4grliancph
- Inference engine: llama.cpp (llama-server)
- Model: orcarouter/Qwen3.8-27B-Uncensored-GGUF — Q4_K_M, 15.7 GiB, with reasoning, tool calling, vision + 0.9 GiB vision projector (mmproj)
- Model storage path: /workspace/models/
- HTTP:
- OpenAI-compatible API (requires authentication)
- llama.cpp web UI (requires authentication)
- Port 8000 exposed via public proxy URL
- SSH: Open with key
- Agentic: built-in web UI with MCP tool support; initial setup has only two benign tools enabled (Current time, Runtime info)
Setup accounts with Runpod and Hugging Face.
Open this model page while logged in to your Hugging Face account and click the request-access / agree button. Wait until it shows as granted.
Navigate to this Runpod template and click 'Configure Pod'.
Click 'Set overrides' and edit the 'Environment variables'. Edit HF_TOKEN to be your Hugging Face API key which you can get from their website. Then add LLAMA_API_KEY with an arbitrary API key value you make up using a random character generator, etc. WARNING: If you don't configure the LLAMA_API_KEY variable then your LLM will be exposed to the public without any authentication!
Set the GPU to be 1 x RTX A6000. Then set the container disk to be 20 GB
and a volume disk of 40 GB which is mounted at /workspace. Sometimes Runpod will be out of capacity for this setup so you may need to play around with other similar configurations that match the template.When it first starts it needs time to download 16.5 GB but later restarts reuse the volume and take under a minute. To monitor the init process get your SSH command from your pod's properties. (You may need to add your SSH key there.) Then SSH and observe with
tail -f /workspace/init.log.After it's up, navigate to your Runpod proxy URL shown in the pod's properties under HTTP services. The web application will ask for your API key to authenticate.
You can also access the API using this:
curl https://<pod-id>-8000.proxy.runpod.net/v1/chat/completions \
-H "Authorization: Bearer <api-key>" \
-H "Content-Type: application/json" \
-d '{"model":"Qwen3.8-27B-Uncensored","messages":[{"role":"user","content":"Hello!"}]}'
Out of the box this has limited agentic support involving only the minimal 'current time' and 'runtime info' tools. It is possible to setup more agent tools but there are major security concerns with setting up an uncensored AI with agentic support that has Internet access. (Reference the OpenAI–HuggingFace incident.)
It goes without saying that this setup is very powerful and should only be used for ethical purposes. Also, it's important to consider the ramifications of any client data being placed in prompts.
Enjoy!




