Skip to main content

Command Palette

Search for a command to run...

Deploying an Uncensored Cloud LLM

Updated
•4 min read•View as Markdown
Deploying an Uncensored Cloud LLM
J
I grew up in the (TRS-)80's playing and creating text based adventure games on monochrome displays. My first infosec experience was booting a hacker from a dial-up BBS at around 12 years old. I forsook technology for a while and became a professional violinist, but in 2006 I decided to launch full career in computers because the money is better! I've been hacking professionally since 2015 and love riding the waves of constant change. I have done well for myself by not expecting others to teach me and just creating my own test environments from scratch. Now I have a successful role at a pentesting firm where I pentest Big 5 clients and play electric violin when I feel like it.

It has long been accepted in the infosec community that open source exploits do more to help defenders than to hurt. If only malicious actors have exploits and tooling then vulnerabilities cannot be detected with precision and actionable results. The world of AI/LLM has generally not caught up to this. Legitimate penetration testers are often fighting with LLMs that flag their security testing prompts as malicious and hamper operations.

I myself have grown tired of trying to use clever prompts to convince LLMs to assist in Red Team exploits. For that reason, I finally found time to spin up my own open source uncensored/abliterated LLM in the cloud that is generally affordable and reasonably easy to setup. This LLM answers easily to prompts regarding technical assistance with penetration testing. Below are my specs and steps for setup.

- Hosting: runpod.io 'secure cloud' container
- Cost: $0.54/hr total
- GPU: RTX A6000 (48 GB VRAM)
- vCPU: 16 vCPUs on an AMD EPYC 7543
- RAM: 62 GB
- Container disk: 20 GB
- Volume disk: 40 GB, mounted at /workspace (18 GB used after initial build)
- GPU driver: 595.91.07
- CUDA version: 13.2
- Container image: ghcr.io/ggml-org/llama.cpp:server-cuda
- Template: Qwen3.8-27B-Uncensored-Q4_K_M (24GB VRAM, 192k context, OpenAI API, llama.cpp) — Template ID 4grliancph
- Inference engine: llama.cpp (llama-server)
- Model: orcarouter/Qwen3.8-27B-Uncensored-GGUF — Q4_K_M, 15.7 GiB, with reasoning, tool calling, vision + 0.9 GiB vision projector (mmproj)
- Model storage path: /workspace/models/
- HTTP:
    - OpenAI-compatible API (requires authentication)
    - llama.cpp web UI (requires authentication)
    - Port 8000 exposed via public proxy URL
- SSH: Open with key
- Agentic: built-in web UI with MCP tool support; initial setup has only two benign tools enabled (Current time, Runtime info)
  1. Setup accounts with Runpod and Hugging Face.

  2. Open this model page while logged in to your Hugging Face account and click the request-access / agree button. Wait until it shows as granted.

  3. Navigate to this Runpod template and click 'Configure Pod'.

  4. Click 'Set overrides' and edit the 'Environment variables'. Edit HF_TOKEN to be your Hugging Face API key which you can get from their website. Then add LLAMA_API_KEY with an arbitrary API key value you make up using a random character generator, etc. WARNING: If you don't configure the LLAMA_API_KEY variable then your LLM will be exposed to the public without any authentication!

  5. Set the GPU to be 1 x RTX A6000. Then set the container disk to be 20 GB
    and a volume disk of 40 GB which is mounted at /workspace. Sometimes Runpod will be out of capacity for this setup so you may need to play around with other similar configurations that match the template.

  6. When it first starts it needs time to download 16.5 GB but later restarts reuse the volume and take under a minute. To monitor the init process get your SSH command from your pod's properties. (You may need to add your SSH key there.) Then SSH and observe with tail -f /workspace/init.log.

  7. After it's up, navigate to your Runpod proxy URL shown in the pod's properties under HTTP services. The web application will ask for your API key to authenticate.

  8. You can also access the API using this:

curl https://<pod-id>-8000.proxy.runpod.net/v1/chat/completions \
  -H "Authorization: Bearer <api-key>" \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen3.8-27B-Uncensored","messages":[{"role":"user","content":"Hello!"}]}'

Out of the box this has limited agentic support involving only the minimal 'current time' and 'runtime info' tools. It is possible to setup more agent tools but there are major security concerns with setting up an uncensored AI with agentic support that has Internet access. (Reference the OpenAI–HuggingFace incident.)

It goes without saying that this setup is very powerful and should only be used for ethical purposes. Also, it's important to consider the ramifications of any client data being placed in prompts.

Enjoy!

20 views