CloudPC.live — Model API
Works with any OpenAI-compatible client

Call your models directly from code

Every AI model on CloudPC.live runs behind an OpenAI-compatible chat API (plus a native Ollama API), so you can call it from a script, a backend service, or any HTTP client — no browser required.

Overview

Your instance's Base URL is the same link you use for Open Web Viewer — it's a full web app (Open WebUI) with a real HTTP API underneath.

Base URLhttps://your-instance.app.cloudpc.live
AuthAuthorization: Bearer <token>
Chat formatOpenAI Chat Completions — POST /api/chat/completions
Native formatOllama API — POST /ollama/api/generate
Content typeapplication/json

Your instance only needs to be assigned and running — open it once from CloudPC.live (Web Viewer or the app itself) so it's warm, then everything below works the same whether a human or your code is calling it.

1 Authentication

Sign in with your instance login (email + password) to get a bearer token — use it in every request below.

Sign in to get a token

Use the same email/password shown in Instance Configuration (crAutoAdmin instances) or your own Open WebUI login.

POST /api/v1/auths/signin
curl -s -X POST "BASE_URL/api/v1/auths/signin" \
  -H "Content-Type: application/json" \
  -d '{"email":"you@example.com","password":"your-password"}'

The response's token field is your bearer token for the calls below.

The token has a long expiry, but it's still a login session, not a permanent credential — if it ever stops working, just sign in again to get a fresh one.

API keys

Open WebUI normally also supports long-lived API keys (Settings → Account → API Keys), but key creation is currently turned off platform-wide on CloudPC instances — use the sign-in token above instead for now.

2 List available models

See exactly which model ids are live on this instance right now, and whether they're loaded into GPU memory.

GET /api/models
curl -s "BASE_URL/api/models" \
  -H "Authorization: Bearer YOUR_TOKEN"
Use the id values from this response as the model field in Chat Completions — they match what you see in the model picker in the app.
Each instance runs exactly one model — whichever one was selected when it was started. Picking a different model in the app's dropdown while you're already connected does not switch the running instance to it; it only takes effect the next time you start a fresh instance. Calling the API with any other model id — including one from the vision-support table below — fails with {"detail":"Model not found"}, verified live. Always call GET /api/models first and use the id it actually returns, rather than assuming a model from the catalog is loaded.

3 Chat completions

Standard OpenAI-shaped request and response — drop-in compatible with most OpenAI client libraries.

POST /api/chat/completions
curl -s -X POST "BASE_URL/api/chat/completions" \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MODEL_ID",
    "messages": [
      {"role": "user", "content": "Give me a one-sentence summary of what you can help with."}
    ]
  }'
import requests

resp = requests.post(
    "BASE_URL/api/chat/completions",
    headers={"Authorization": "Bearer YOUR_TOKEN"},
    json={
        "model": "MODEL_ID",
        "messages": [
            {"role": "user", "content": "Give me a one-sentence summary of what you can help with."}
        ],
    },
)
print(resp.json()["choices"][0]["message"]["content"])
const resp = await fetch("BASE_URL/api/chat/completions", {
  method: "POST",
  headers: {
    "Authorization": "Bearer YOUR_TOKEN",
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "MODEL_ID",
    messages: [
      { role: "user", content: "Give me a one-sentence summary of what you can help with." },
    ],
  }),
});
const data = await resp.json();
console.log(data.choices[0].message.content);

4 Sending an image

Only some models can see images — check the table below, or check GET /api/models for a model's family before assuming.

This table lists what's vision-capable in the catalog — it does not mean that model is loaded on your instance right now. Your instance only runs the one model it was started with (see the note in step 2). To actually use one of the vision models below, start a new instance and pick it there first.

Which models currently support image input

Supports imagesModel idDisplay name
✅ Yesgemma4:12bGemma 4 (12B) — Lightweight Medical Vision & Text QA
✅ Yeshf.co/Mungert/Lingshu-32B-GGUF:Q3_K_MLingshu 32B — Medical Vision-Language QA
✅ Yes — verified livehf.co/unsloth/medgemma-1.5-4b-it-GGUF:Q8_0MedGemma 1.5 (4B) — CT/MRI & Pathology (Official)
✅ Yes — verified liveqwen3.8:27bQwen 3.8 (27B) — Multimodal Charts, EKGs & Clinical Reports
✅ Yes — verified liveminicpm-v4.5:q8_0MiniCPM-V 4.5 (8B) — OCR, Documents & High-FPS Video
✅ Yes — verified livenemotron3:33b-q8NVIDIA Nemotron 3 Nano Omni (33B) — Documents, Charts & Screenshots
✅ Yes — verified livecosmos-reason2:8b-q8_0NVIDIA Cosmos Reason 2 (8B) — Physical-World & Spatial Reasoning
✅ Yes — verified livemuse-glimmer:30b-q8_0Meta Muse Glimmer (30B) — Agentic Tasks, Code & Screenshots
❌ Text only — verified livealibayram/medgemma:4bMedGemma 4B — Clinical QA
❌ Text onlymedgemma:27bMedGemma 27B — Clinical Reports & Diagnostics
❌ Text onlygemma4:26bGemma 4 (26B) — Clinical Report Analysis (Text-Only)
❌ Text onlyqwen3.6:27bQwen 3.6 (27B) — Clinical Data Reasoning (Text-Only)
❌ Text onlydeepseek-r1:14bDeepSeek R1 (14B) — Clinical Reasoning
❌ Text onlyopenbiollm-70b-chat:Q4_K_MOpenBioLLM 70B — Comprehensive Biomedical Reasoning
❌ Text onlyhf.co/bartowski/baichuan-inc_Baichuan-M2-32B-GGUF:Q6_KBaichuan M2 32B — Medical Reasoning (HealthBench)
❌ Text onlygpt-oss:120bGPT-OSS 120B — General Medical Reasoning (OpenAI)
Display names now flag text-only vs. vision explicitly, so what you see in the app's model picker matches this table. This list reflects the current catalog and may change as models are added.

Request format

Same endpoint as a normal chat, but content becomes an array with a text part and an image_url part (a data URI, or a public image URL) — the standard OpenAI vision format.

POST /api/chat/completions
curl -s -X POST "BASE_URL/api/chat/completions" \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MODEL_ID",
    "messages": [
      {
        "role": "user",
        "content": [
          {"type": "text", "text": "What do you see in this image?"},
          {"type": "image_url", "image_url": {"url": "data:image/png;base64,BASE64_IMAGE_DATA"}}
        ]
      }
    ]
  }'

Verified live: a test image (a blue square, top-left, and a red circle, bottom-right) sent to medgemma-1.5-4b got back — correctly — "I see a blue square and a red circle in the image. The blue square is located in the top left position. The red circle is located in the bottom right position."

Sending an image to a text-only model

You'll get a clean 400, not a silent failure — confirmed live against alibayram/medgemma:4b:

HTTP 400
{"detail":"{\"error\":{\"code\":400,\"message\":\"Multimodal data provided, but model does not support multimodal requests.\",\"type\":\"invalid_request_error\"}}"}

5 Streaming responses

Set "stream": true to get tokens as they're generated, as server-sent events — same shape OpenAI's streaming API uses.

curl -N -X POST "BASE_URL/api/chat/completions" \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MODEL_ID",
    "messages": [{"role": "user", "content": "Count from 1 to 5."}],
    "stream": true
  }'

Each line arrives as data: {json chunk}, ending with data: [DONE] — parse it the same way you'd parse an OpenAI streaming response.

Using an OpenAI client library

Because the API is OpenAI-compatible, you can often just point an existing OpenAI SDK at your instance instead of hand-rolling requests.

from openai import OpenAI

client = OpenAI(
    base_url="BASE_URL/api",
    api_key="YOUR_TOKEN",
)

resp = client.chat.completions.create(
    model="MODEL_ID",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "BASE_URL/api",
  apiKey: "YOUR_TOKEN",
});

const resp = await client.chat.completions.create({
  model: "MODEL_ID",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(resp.choices[0].message.content);
Also want the raw Ollama-native API instead of the OpenAI shape? Same instance, different path: POST BASE_URL/ollama/api/generate and GET BASE_URL/ollama/api/tags.

Notes & troubleshooting

Your instance has to be running

If it's paused or hasn't been opened yet, requests will fail or time out. Start it once from CloudPC.live (or hit any endpoint here — a paused instance won't wake up for API calls the way it does for the web viewer).

401 Unauthorized

Your token expired or was revoked. Sign in again (Authentication) to get a fresh one.

{"detail":"Model not found"}

The most common cause: the model id in your request isn't the one actually loaded on this instance — each instance runs exactly one model at a time (see the note in step 2). Confirmed live: picking a different model in the app's dropdown, or copying a model id from the vision table above, does not change what's running until a new instance is started with that model selected.

Check GET /api/models first — the id you send has to match what's listed there ("loaded": true means it's warm and ready; otherwise the first request may take longer while it loads).

CORS / calling from a browser app

These endpoints are meant to be called from a server, script, or backend you control — not directly from another website's front-end JavaScript with your token embedded in it.