PIES Studio · Operations Manual
Run PIES Ultra — the AI engine behind PIES Studio — entirely on your own GPU infrastructure. Your prompts, application definitions, and data never leave your network. This manual covers installation, connecting PIES Studio, day-to-day operation, and troubleshooting.
PIES Private AI is a single Docker container that serves the PIES Ultra v0.9 model — the model PIES ships and certifies (17B active / 400B total parameters) — through an OpenAI-compatible API, plus an administration API that PIES Studio manages it through. You can also run models of your own alongside it (section 9). It runs only with a valid PIES license and stores everything — model weights, your license, API keys — in one Docker volume on your server.
| Requirement | Detail |
|---|---|
| GPU | NVIDIA, 256 GB+ total VRAM — e.g. 4× A100 80GB, 2× H200, 2× B200. Weights are 216 GB; the rest is KV cache and headroom, so 256 GB is the floor, not the target. Verified on 2× H200 and 2× B200. For evaluation, PIES Ultra Lite needs one 94 GB+ GPU (H100 NVL or H200), and an Apple Silicon Mac with 96 GB+ unified memory runs it for small teams — see section 3 |
| OS | Linux x86_64 |
| Software | Docker with the NVIDIA Container Toolkit |
| Disk | ~350 GB free — 216 GB of weights plus download temp space |
| Network | Outbound HTTPS for the one-time weights download; port 8000 reachable by PIES Studio |
| License | Your pies_studio.license file, provided by your PIES account manager |
No license, no inference. The service starts without a license but refuses all AI requests until a valid, unexpired license is installed. Health and administration endpoints stay available so you can always see why.
Any machine meeting the specification above works. Four setups worth naming, in the order most organisations consider them: a server you own, a GPU rented by the hour while you evaluate, a cloud VM inside your own subscription, and — for smaller teams — an Apple Silicon Mac.
PIES Private AI was built for this. The model runs on hardware you own, in a building you control, and nothing about a request — the prompt, your data, your application — leaves the room. It is the only option that can be made genuinely air-gapped, and the only one where the running cost is the electricity.
| What you need | Detail |
|---|---|
| A GPU server with 256 GB+ VRAM | 4× A100 80GB, 2× H200 or 2× B200 all clear it. An NVIDIA DGX or any vendor's equivalent works — nothing here is specific to a brand of chassis |
| Linux x86_64, ~350 GB free | 216 GB of weights plus room to download them. Local NVMe keeps load times short |
| NVIDIA drivers + Container Toolkit | The only host software required. Everything else ships inside the container |
| Reachable on port 8000 | From your PIES Studio hosts only. No inbound internet, no public address |
Internet access is needed once, to download the weights — and not even then if you bring them in on media, which is covered in section 7. After that the server can be disconnected entirely and the platform keeps working: the license is validated locally, not by calling home.
Buying advice, briefly. Fewer, larger cards beat more small ones: the weights are split across GPUs, so 2× 141 GB is simpler and faster than 8× 40 GB. Total VRAM is what matters, and 256 GB is the floor rather than the target — headroom above it becomes KV cache, which is what lets several people build at once.
The cheapest way to stand PIES Private AI up, and the one to use while you are evaluating. Providers such as RunPod, Lambda or Vast rent GPU machines by the hour and let you stop them when you are not using them.
| Choose | Why |
|---|---|
| 2× H200 or 2× B200 | 256 GB+ VRAM in the fewest cards. Both verified by PIES |
| A persistent network volume of 550 GB+ | This is the important one. Weights live on the volume, so stopping the machine keeps them — a restart takes minutes rather than re-downloading 216 GB |
| The container port 8000 exposed | PIES Studio talks to it here. Most providers give you a proxied HTTPS URL |
How you install depends on what the provider hands you. A rented virtual machine is an ordinary Linux host with GPUs: install exactly as in section 4. A rented GPU pod (RunPod, Vast and similar) is itself a container, with no Docker inside it, so the installer cannot run there. Use the pod setup below instead. What differs from your own server is the running cost: you pay per hour while it runs, and only for the volume while it is stopped. That makes it practical to run the AI during working hours and stop it overnight; the model reloads into VRAM in 10–20 minutes on the next start.
Check the machine you get back. After a stop/start some
providers reallocate hardware, and a machine that comes back with fewer GPUs
than it had will fail to load the model — with an out-of-memory error rather
than an obvious explanation. Confirm the GPU count after every start:
nvidia-smi -L.
On a pod you do not run pies-llm install. You give the
provider the PIES image and a few settings, and the pod runs it directly.
Everything the installer would have done is in this table.
| Setting | Value |
|---|---|
| Container image | ghcr.io/pies-io/pies-llm:vllm-v0.9 (pin the version) |
| GPUs | 2× H200 or 2× B200 |
| Volume | 550 GB+, mounted at /app/models |
| Exposed port | 8000, as HTTP. The provider gives you an HTTPS address for it |
PIES_LLM_MODEL | slab (PIES Ultra) |
PIES_LLM_API_KEY | Your admin key. Generate one, for example echo pies-$(openssl rand -hex 24), and keep it — Studio needs it |
PIES_LLM_N_CTX | 65536, the same context window the installer sets |
PIES_LLM_LICENSE_B64 | Your license file, base64-encoded: base64 -i pies_studio.license | tr -d '\n' |
On first start the pod writes the license onto the volume, downloads the weights to it, and loads the engine. Because both live on the volume, later starts reuse them. Then connect PIES Studio as in section 4, using the provider's HTTPS address for port 8000 and your admin key.
The license travels as a pod setting. Anyone with access
to your provider account can read pod settings. Once the pod has started
once, the license is on the volume and you can remove
PIES_LLM_LICENSE_B64 from the pod.
PIES_LLM_LICENSE_B64 is read only when the volume holds no
license yet (/app/models/pies_studio.license). It never replaces
one that is already there, so changing the pod setting does not renew a
license. To renew on a pod, open the pod's terminal and replace
/app/models/pies_studio.license with the new file; it is
re-checked within a minute, no restart needed.
For evaluation, or to check networking, the license and the Studio connection before renting 256 GB of VRAM, run PIES Ultra Lite instead. It is a smaller model (~65 GB of weights) that fits on one GPU and is ready about 2½ minutes after the pod is created. It answers fast but builds noticeably less well than PIES Ultra, so use it to try the platform, not for real work. It uses the llama.cpp image, on a pod or a server alike:
| Setting | Value |
|---|---|
| Container image | ghcr.io/pies-io/pies-llm:cuda-v0.9 |
| GPU | One NVIDIA card with 94 GB+ — H100 NVL (94 GB) or H200 (141 GB). Not a Blackwell card (B200, B300, RTX PRO 6000): this image is built for CUDA 12.4, which cannot run on them |
| Volume | 120 GB+, mounted at /app/models |
PIES_LLM_MODEL | scout (PIES Ultra Lite) |
PIES_LLM_BACKEND | gguf |
PIES_LLM_N_GPU_LAYERS | -1 (the whole model on the GPU) |
PIES_LLM_N_CTX | 131072. A 32k window is too small for a Studio build |
PIES_LLM_API_KEY, PIES_LLM_LICENSE_B64 | As for PIES Ultra above |
Studio shows it as PIES Ultra Lite (dev). Everything else — connecting Studio, keys, the license — works exactly as for PIES Ultra.
What most organisations move to for production: the machine lives inside your tenancy, your network controls, and your compliance boundary. On Azure the relevant families are ND A100 v4 (8× A100 80 GB) and ND H100 v5 (8× H100 80 GB); AWS and Google have direct equivalents.
| Step | Detail |
|---|---|
| Size the VM | Any GPU VM totalling 256 GB+ VRAM. The 8-card families exceed it comfortably, leaving room for a large KV cache |
| Attach a data disk | 512 GB+ Premium SSD for the weights, mounted where the container's volume lives. Keep it separate from the OS disk so the VM can be resized or rebuilt without another download |
| Install GPU drivers | The provider's GPU driver extension, then the NVIDIA Container Toolkit so Docker can see the cards |
| Restrict the port | Allow 8000 only from your PIES Studio hosts, in the network security group. The service is licensed and key-protected, but it should not be open to the internet |
| Keep it private | No public IP is needed if Studio reaches it over your virtual network or a peering |
Reserved or committed-use pricing changes the economics substantially for a machine that runs continuously — worth pricing before choosing between hourly rental and a cloud VM.
On cost. GPU pricing moves quickly and differs by region and commitment, so we do not quote figures that would be stale by the time you read them. The shape holds, though: renting by the hour is far cheaper while you are evaluating and stopping the machine between sessions, and a committed cloud VM wins once the AI is in constant use. The one cost that persists either way is storage for the weights.
A Mac with Apple Silicon runs PIES Ultra Lite on the same llama.cpp engine as the Docker image, using Apple's Metal GPU. It serves one request at a time and answers more slowly than a multi-GPU server, which makes it a fit for a small team or an evaluation on hardware you may already own — not for many people building at once. PIES Ultra itself belongs on a GPU server: on a Mac the launcher runs Lite and refuses PIES Ultra unless you override it deliberately (below).
| What you need | Detail |
|---|---|
| An Apple Silicon Mac, 96 GB+ unified memory | The model is ~65 GB; the rest keeps the machine usable while it is loaded |
| ~100 GB free disk | The one-time weights download (~65 GB) plus headroom |
| macOS with Xcode Command Line Tools | xcode-select --install — needed once, to build the engine |
Install with one command in Terminal — it downloads PIES Private AI to
~/pies-private-ai:
curl -fsSL https://github.com/pies-io/pies-studio-releases/releases/download/private-ai-mac-v0.9.1/install-mac.sh | bash
Then, from that folder, three commands finish the job:
cd ~/pies-private-ai
./pies-llm.sh setup # one-time: engine + model download
./pies-llm.sh license /path/to/pies_studio.license
./pies-llm.sh start # loads the model, waits until ready
The same license rules apply as on a server: without a valid
pies_studio.license the service starts but refuses AI requests.
The model loads in about two minutes and appears in Studio as PIES Ultra Lite (dev). It builds less well than PIES Ultra on a GPU server; use the Mac for evaluation and small teams.
PIES Ultra on a Mac is possible only as a deliberate override, on a Mac with 256 GB+ unified memory and ~250 GB free disk (a ~200 GB download), with nothing else loaded first:
./pies-llm.sh setup --ultra
PIES_LLM_ALLOW_ULTRA_MAC=1 ./pies-llm.sh start --ultra
All commands are safe to re-run — setup skips what is already done and a
stopped download resumes where it left off. ./pies-llm.sh status
shows health, and PIES Studio connects to port 8000 exactly as with the
Docker install (section 4, step 4).
Run the installer again with --upgrade. It replaces the program
only — your models, training data, adapters, API keys and license are left
exactly as they are, so no weights are downloaded twice:
curl -fsSL https://github.com/pies-io/pies-studio-releases/releases/download/private-ai-mac-v0.9.1/install-mac.sh | bash -s -- --upgrade
cd ~/pies-private-ai && ./pies-llm.sh restart
Without --upgrade the installer refuses to touch an existing
folder, so re-running the plain install line can never overwrite a working
service. The version now serving is shown in
Private LLM → Manage, which is the quickest way to confirm an
upgrade took.
On an Apple Silicon Mac, skip this section. Docker is not used
on a Mac — macOS gives containers no GPU access, so the
docker run commands below fail there by design. Use the Apple Silicon installer in section 3 instead; everything from
section 5 onward (connecting PIES Studio) applies to both.
Run everything below on the GPU server, as a user who can
run docker. Four steps; the installer does the checking.
pies-llm commandThe management CLI ships inside the container image — no separate download:
docker run --rm --entrypoint cat ghcr.io/pies-io/pies-llm:vllm /app/pies-llm-ctl.sh \
| sudo tee /usr/local/bin/pies-llm >/dev/null && sudo chmod +x /usr/local/bin/pies-llm
pies-llm install /path/to/pies_studio.license
It checks your GPUs, free disk and Docker GPU access, pulls the image, stores the license, generates your admin key and starts the service. If a prerequisite is missing it stops and says what to fix — nothing large is downloaded until the checks pass.
The installer prints your admin API key —
save it. It is also kept in ~/.pies-llm/config.
pies-llm logs -f
First start downloads 216 GB of model weights into the
pies-llm-models volume, then loads them into VRAM. The
service is deliberately quiet until the engine is ready — a port that
does not answer yet is normal, not a hang.
| Stage | Typical time |
|---|---|
| Weights download (once, cached on the volume) | 30–90 min |
| Engine load into VRAM (every start) | 10–20 min |
| Later starts (weights already present) | 10–20 min total |
pies-llm status reports model_loaded: false
until the engine finishes; then it shows backend: vllm and
your licence details.
pies-llm status # model_loaded: true, backend: vllm, license valid
Open port 8000 to PIES Studio, then in Studio open
PIES AI → AI Settings (an administrator's view) and, under
Providers, find the PIES Private AI
row and enter:
http://<your-server>:8000
(on a GPU pod, the provider's HTTPS address for port 8000)Press Connect, then Test connection,
and check the Private LLM tab shows the model Online.
The URL and key apply to the whole install. Run the CLI on the server
itself — it talks to the service on localhost.
Connecting makes the model available; which requests actually use it is decided by your AI governance policies — see section 10.
| Command | What it does |
|---|---|
pies-llm status | Health, model state, license validity (with reason), GPUs |
pies-llm start / stop / restart | Control the service |
pies-llm logs -f | Follow live logs |
pies-llm debug | One diagnostic bundle to send to PIES support |
pies-llm license show | Organization, tier, expiry, days remaining |
pies-llm license install <file> | Install or renew the license (validated first, no restart) |
pies-llm key create <name> | Mint an inference-only API key (shown once) |
pies-llm key list / key delete <id> | List (masked) or revoke keys |
pies-llm uninstall | Remove the service; weights and license are kept |
Most day-to-day operation no longer needs a terminal. PIES AI → Private LLM → Manage works against whichever host Studio is connected to and shows:
| What you see | What you can do |
|---|---|
| This host — the pies-llm build version, engine (llama.cpp or vLLM), device, and license state | Reload the loaded model; confirm which build is actually serving after an upgrade |
| Adapters — every fine-tuned adapter on the host, and which one is serving | Promote an adapter, or step back to the base model |
| API keys — inference keys, masked | Mint a key (shown once) or revoke one |
| Service log — what the service reported while starting and loading | The first place to look when a model will not load |
Starting and stopping the service itself stays on the host, for the obvious
reason: when it is down there is nothing for Studio to talk to. Use
pies-llm start there.
The admin key (created at install) opens everything and is
what PIES Studio uses. Inference keys
(pies-llm key create) can only run AI requests — hand them to
individual apps or teams, and revoke any one of them at any time without
affecting the others.
The Mac install manages keys with the same commands, through
./pies-llm.sh. The admin key is created
automatically the first time you run setup or
start — it is printed once and kept in
~/.pies/pies-llm.env:
./pies-llm.sh key show # the admin key — what PIES Studio connects with
./pies-llm.sh key create <name> # mint an inference-only key (shown once)
./pies-llm.sh key list # list minted keys (masked)
./pies-llm.sh key delete <id> # revoke one
Admin vs. inference works exactly as on a server: the admin key opens everything and belongs in PIES Studio's AI Settings; inference keys can only run AI requests and can be revoked individually at any time.
PIES issues you a single file, pies_studio.license — either
attached to an email or downloaded from the setup link your account manager
sends. Copy it to the GPU server and point the installer at it. It is stored
on the pies-llm-models volume beside the model weights, so it
survives restarts, container recreation and upgrades; the copy you were sent
is not needed afterwards.
pies-llm license show # org, tier, expiry, days remaining
pies-llm license install ~/pies_studio.license # install or renew
Renewal needs no restart — the file is re-checked within a
minute. The new license is validated before it replaces the old one, so a
wrong or corrupt file is rejected and your running license is left untouched.
pies-llm license show warns when 30 days or fewer remain.
The license is checked continuously — signature, expiry, and status. When it expires or is removed, AI requests return HTTP 403 with the exact reason, while health and administration stay reachable so you can always see why.
The only step that needs the internet is the one-time weights download. Do it on any connected machine, carry the result across, and import it.
Get the weights (216 GB). This is a public Hugging Face repository — no
PIES account or token is involved. Either route must end with a folder named
exactly Llama-4-Maverick-17B-128E-Instruct-quantized.w4a16.
Option A — command line. Resumable, and strongly preferred at this size: an interrupted download continues where it left off, just run it again.
pip install huggingface_hub
huggingface-cli download RedHatAI/Llama-4-Maverick-17B-128E-Instruct-quantized.w4a16 \
--local-dir Llama-4-Maverick-17B-128E-Instruct-quantized.w4a16
Option B — manual browser download, if the machine has no
Python or your policy requires a human-reviewed download. Open the
repository
file list, create a folder named
Llama-4-Maverick-17B-128E-Instruct-quantized.w4a16, and save
every file into it:
| File(s) | Notes |
|---|---|
model-*.safetensors | ~50 shards — the bulk of the 216 GB |
model.safetensors.index.json | Lists the shards; the import check reads this |
config.json, generation_config.json | Model configuration |
tokenizer.json, tokenizer_config.json, special_tokens_map.json | Tokenizer |
Keep the structure flat — no sub-folders, no renaming. Browser downloads have no resume and no integrity check, which is why Option A is preferred.
Either way, copy the whole directory onto your approved
media — a missing shard is the usual air-gap failure, and
import-weights checks for exactly that on the other side before
it copies anything. Take the container image the same way:
docker pull ghcr.io/pies-io/pies-llm:vllm
docker save ghcr.io/pies-io/pies-llm:vllm | gzip > pies-llm-vllm.tar.gz
docker load < pies-llm-vllm.tar.gz
pies-llm import-weights /media/transfer/Llama-4-Maverick-17B-128E-Instruct-quantized.w4a16
pies-llm start # finds the weights on the volume — never touches the network
import-weights checks the directory is a complete snapshot
before copying 216 GB, and stores it on the
pies-llm-models volume under the exact name the service looks
for — so it works even if your media folder is named something else.
After this, nothing in the product requires egress — prompts, data, and license checks all stay on the server.
Out of the box the service speaks plain HTTP on port 8000 — appropriate on localhost, behind a TLS-terminating proxy (nginx, caddy, a cloud load balancer), or on an isolated network. If the API key crosses a network you don't fully trust, serve HTTPS directly by mounting a PEM certificate and key:
docker run ... \
-v /etc/pies-llm/certs:/certs:ro \
-e PIES_LLM_SSL_CERT=/certs/fullchain.pem \
-e PIES_LLM_SSL_KEY=/certs/privkey.pem \
ghcr.io/pies-io/pies-llm:vllm
Then use https://<server>:8000 in PIES Studio.
PIES Private AI runs models beyond PIES Ultra. The Private LLM tab has an Upload Your Own tile: bring any GGUF model — from Hugging Face or your own fine-tune — and the platform handles the rest, with no configuration.
| Step | What happens |
|---|---|
| Check | Before the upload starts, the host checks it could actually serve a model that size — being told after a two-hour transfer that the box is too small is the worst possible moment |
| Upload | The file is streamed to disk beside the built-ins and appears as a card. Uploading under a name that already exists offers Replace; the model currently serving cannot be replaced underneath itself |
| Understand | The service reads the model's own capabilities — context window, size, family — directly from the file; nothing to register |
| Speak PIES | Tool calls from uploaded models are grammar-constrained to the exact shape the platform expects, whatever convention the model was trained on |
| Certify | Test model makes it complete a small real build, graded by the server — the model never grades itself. The verdict lands on the card: build-capable, chat-only, or failed |
| Govern | Only models certified build-capable are given builds; chat-only models still answer questions. A weak upload degrades gracefully instead of producing broken apps |
Uploaded models also join the fine-tuning loop: verified wins from their builds are collected, and an adapter is only promoted to serving if it beats the model's own certified baseline.
Seed training data ships with every install. The file
data/collected/seed_mcp_pack.jsonl is built from the release's PIES Studio
MCP training pack — the canonical widget, step, event and process shapes — and
stamped with the pack version and hash. It is copied into the training data
directory on first start only if no seed file is there; your own training files
are never overwritten or removed. Its examples are counted in the training data totals,
and every fine-tuning job trains on it together with your own data, so even a
fresh install has a starting corpus. To train without it, send
"include_seed": false with the training request; if the file is
missing, training carries on without it. To
take a newer release's seed, delete the file and restart.
Certify before you rely on it. An uncertified model is not offered builds, and a model that has never been graded is an unknown quantity — smaller models are noticeably weaker at calling tools correctly. Run Test model once after uploading, and again after promoting an adapter.
Once connected, PIES Private AI is one of the providers Studio can use. Your AI governance policies decide which requests go to it — the full rules are in the user manual's AI governance chapter. The parts that matter for Private AI:
| Topic | What happens |
|---|---|
| Auto (governed) or Force one model | Under AI Settings, Auto lets the policies route each request; its Can use line lists PIES Private AI only while Studio can reach it. Force one model pins every request to one model you pick — redaction and audit still apply |
| Default Balanced | The policy every organisation starts with. It allows cloud models and PIES Private AI, prefers cloud models, and names GPT-4o as its fallback. If a request it sends to Private AI gets no answer, the fallback model answers instead and the audit row records fallback_allowed |
| Air-Gapped | Sends every request — builds, chat, agents, and the assistants inside your deployed applications — to PIES Private AI, and blocks cloud providers outright. There is no cloud fallback |
| Compliance packs (HIPAA and others) and Regulated Industry | Send only the sensitive classes to Private AI. The HIPAA pack keeps PHI on Private AI and refuses the request when Private AI is unavailable rather than sending it to the cloud; PII may go to a cloud model after masking |
| Which model answers | On the private route the model actually loaded on your host answers, whatever name the policy stores — PIES Ultra, PIES Ultra Lite, or a model you uploaded. The audit log names the model that served each request |
| Builds and agents | Builds and agent runs, including their tool calls (reading tables, saving screens), run on Private AI like any other request. A model drives a build only if it is certified build-capable (section 9). PIES Ultra Lite completes builds, but simpler and slower ones than PIES Ultra or a cloud model — expect to correct more by hand |
| Assistants in deployed apps | An application's own assistant is governed exactly like a build: its conversations follow the same policies, so an air-gapped policy keeps your end users' conversations on Private AI too |
If the server is stopped, still starting, or the URL in AI Settings is wrong, a request routed to Private AI ends with a plain message rather than a connection error — "PIES Private AI is not reachable at the address in AI Settings…" when nothing answers, or "PIES Private AI is not available: the requested model … did not answer…" when a proxy answers but no model is serving. What happens next depends on the policy:
Agent runs fail fast with the same message on their run record. On a GPU pod, remember the server stops when you stop the pod: an air-gapped Studio pointed at a stopped pod cannot run any AI until the pod is started again.
| Symptom | Cause & fix |
|---|---|
| Studio shows Offline right after install | The one-time weights download (216 GB) and the 10–20 minute engine load are still running — watch pies-llm logs -f. The service answers only once the engine is ready. |
| The port does not answer at all | Normal while the weights download: the service starts only after that. Follow the download in pies-llm logs -f, or in your provider's log view on a GPU pod. Once the service is up, the whole boot log, including earlier boots, is at /v1/admin/boot-log with your admin key, and pies-llm debug includes it. |
| On a GPU pod, the provider's address returns 404, then 502 | Both are normal stages. 404 means the container has not started yet, usually while the image is pulled. 502 means it has started but the engine is still loading. Wait for /health to answer. |
Docker is required or Docker daemon not reachable on a rented GPU |
You are on a GPU pod, which is itself a container. Do not run the installer there. Configure the pod directly, as in section 3. |
| A GPU pod cannot be created: "no instances currently available" | The provider has no free cards of that type right now, even if its list shows some in stock. Pick another supported GPU type or region, or try again later. |
The :cuda image will not start on B200 or other Blackwell GPUs |
The llama.cpp image is built for CUDA 12.4, which predates Blackwell. Use the :vllm image, the production image, on these cards. The :vllm image carries PIES Ultra only, so PIES Ultra Lite needs a Hopper or Ampere card instead (H100 NVL or H200, section 3). |
A build fails with Requested tokens (N) exceed context window of M |
The service was started with a context window smaller than a Studio build needs. Raise PIES_LLM_N_CTX (the installer sets 65536 for PIES Ultra; use 131072 for PIES Ultra Lite; a GPU pod takes it as a pod setting) and restart. A larger window uses more VRAM. |
A build stops with context canceled |
A second message was sent while the first build was still running, and the new one replaced it. Wait for a build to finish before sending the next message. |
| A multi-GPU host loads onto one GPU | Set PIES_LLM_TP=<number of GPUs> in ~/.pies-llm/config and pies-llm restart. Auto-detection covers most hosts. |
| Every AI call returns 403 | License missing, expired, or inactive. pies-llm status shows the exact reason in the license block. Install a valid license (section 6). |
could not select device driver when starting |
NVIDIA Container Toolkit missing on the host — install it, then pies-llm start. |
| On a Mac: load refused — "N GB available, the model needs M GB" — with nothing else running | macOS is holding the memory as file cache (a big download or a previous model load). Flush it with sudo purge and retry ./pies-llm.sh start; a reboot — loading the model before anything else — is the sure fix. The launcher's doctor says the same. |
| Model fails to load / out of memory | Total VRAM below 256 GB, or other processes holding GPU memory. Check nvidia-smi; free the GPUs and pies-llm restart. |
| Key rejected in Studio | Studio needs the admin key, not an inference key. Find it in ~/.pies-llm/config on the server, with ./pies-llm.sh key show on a Mac, or in the pod's PIES_LLM_API_KEY setting on a GPU pod. |
| Studio says PIES Private AI is not reachable or is not available | Studio cannot get an answer from the Service URL in AI Settings: the server or pod is stopped, still loading, or the URL is wrong. Check pies-llm status (or the pod) and the URL, then use Test connection. Under an air-gapped policy nothing falls back to the cloud meanwhile (section 10). |
| Anything else | Run pies-llm debug and send the output to PIES support. |
Mostly the one-time download. On a rented 2× H200 pod, PIES Ultra went from creating the pod to answering in under 30 minutes: about 8 minutes to pull the image, 9 to download the 216 GB of weights, and 8 to load the engine. A restart with the weights already on the volume took about 17 minutes. Downloads on your own network may be slower; plan for the ranges in section 4.
You pay by the hour while the machine runs, and for the volume while it is stopped. A 2× H200 pod cost about US$9 an hour in our test, so a first install plus an afternoon of evaluation is tens of dollars. PIES Ultra Lite on one H200 cost about US$4.60 an hour. Stop the machine when you are not using it. To stop paying for the volume too, delete it; the next start then downloads the weights again.
Yes, with an inference key (section 5). The API is OpenAI-compatible, with two differences to know about before you connect a third-party client:
finish_reason is tool_calls, and the call is
a JSON object in the message content:
{"name": "…", "parameters": {…}}. The tool_calls
field itself is empty. Parse the content."stream": true gets the complete response in one piece.The production image batches requests, so several builds and chats run side by side rather than queueing. In our test, four requests sent at once all finished within 7 seconds, faster than running them one after another.
No, as long as the volume is kept. The weights and the license live on the volume. On a GPU pod, the admin key is a pod setting, so it survives too.
Support: support@pies.io — include the output of
pies-llm debug and your organization name. Diagnostic output
contains no prompts or application data.