Call your models directly from code
Every AI model on CloudPC.live runs behind an OpenAI-compatible chat API (plus a native Ollama API), so you can call it from a script, a backend service, or any HTTP client — no browser required.
Overview
Your instance's Base URL is the same link you use for Open Web Viewer — it's a full web app (Open WebUI) with a real HTTP API underneath.
| Base URL | https://your-instance.app.cloudpc.live |
|---|---|
| Auth | Authorization: Bearer <token> |
| Chat format | OpenAI Chat Completions — POST /api/chat/completions |
| Native format | Ollama API — POST /ollama/api/generate |
| Content type | application/json |
Your instance only needs to be assigned and running — open it once from CloudPC.live (Web Viewer or the app itself) so it's warm, then everything below works the same whether a human or your code is calling it.
1 Authentication
Sign in with your instance login (email + password) to get a bearer token — use it in every request below.
Sign in to get a token
Use the same email/password shown in Instance Configuration (crAutoAdmin instances) or your own Open WebUI login.
curl -s -X POST "BASE_URL/api/v1/auths/signin" \
-H "Content-Type: application/json" \
-d '{"email":"you@example.com","password":"your-password"}'
The response's token field is your bearer token for the calls below.
API keys
Open WebUI normally also supports long-lived API keys (Settings → Account → API Keys), but key creation is currently turned off platform-wide on CloudPC instances — use the sign-in token above instead for now.
2 List available models
See exactly which model ids are live on this instance right now, and whether they're loaded into GPU memory.
curl -s "BASE_URL/api/models" \ -H "Authorization: Bearer YOUR_TOKEN"
id values from this response as the model field in Chat Completions — they match what you see in the model picker in the app.{"detail":"Model not found"}, verified live. Always call GET /api/models first and use the id it actually returns, rather than assuming a model from the catalog is loaded.
3 Chat completions
Standard OpenAI-shaped request and response — drop-in compatible with most OpenAI client libraries.
curl -s -X POST "BASE_URL/api/chat/completions" \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model": "MODEL_ID",
"messages": [
{"role": "user", "content": "Give me a one-sentence summary of what you can help with."}
]
}'
import requests
resp = requests.post(
"BASE_URL/api/chat/completions",
headers={"Authorization": "Bearer YOUR_TOKEN"},
json={
"model": "MODEL_ID",
"messages": [
{"role": "user", "content": "Give me a one-sentence summary of what you can help with."}
],
},
)
print(resp.json()["choices"][0]["message"]["content"])
const resp = await fetch("BASE_URL/api/chat/completions", {
method: "POST",
headers: {
"Authorization": "Bearer YOUR_TOKEN",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "MODEL_ID",
messages: [
{ role: "user", content: "Give me a one-sentence summary of what you can help with." },
],
}),
});
const data = await resp.json();
console.log(data.choices[0].message.content);
4 Sending an image
Only some models can see images — check the table below, or check GET /api/models for a model's family before assuming.
Which models currently support image input
| Supports images | Model id | Display name |
|---|---|---|
| ✅ Yes | gemma4:12b | Gemma 4 (12B) — Lightweight Medical Vision & Text QA |
| ✅ Yes | hf.co/Mungert/Lingshu-32B-GGUF:Q3_K_M | Lingshu 32B — Medical Vision-Language QA |
| ✅ Yes — verified live | hf.co/unsloth/medgemma-1.5-4b-it-GGUF:Q8_0 | MedGemma 1.5 (4B) — CT/MRI & Pathology (Official) |
| ✅ Yes — verified live | qwen3.8:27b | Qwen 3.8 (27B) — Multimodal Charts, EKGs & Clinical Reports |
| ✅ Yes — verified live | minicpm-v4.5:q8_0 | MiniCPM-V 4.5 (8B) — OCR, Documents & High-FPS Video |
| ✅ Yes — verified live | nemotron3:33b-q8 | NVIDIA Nemotron 3 Nano Omni (33B) — Documents, Charts & Screenshots |
| ✅ Yes — verified live | cosmos-reason2:8b-q8_0 | NVIDIA Cosmos Reason 2 (8B) — Physical-World & Spatial Reasoning |
| ✅ Yes — verified live | muse-glimmer:30b-q8_0 | Meta Muse Glimmer (30B) — Agentic Tasks, Code & Screenshots |
| ❌ Text only — verified live | alibayram/medgemma:4b | MedGemma 4B — Clinical QA |
| ❌ Text only | medgemma:27b | MedGemma 27B — Clinical Reports & Diagnostics |
| ❌ Text only | gemma4:26b | Gemma 4 (26B) — Clinical Report Analysis (Text-Only) |
| ❌ Text only | qwen3.6:27b | Qwen 3.6 (27B) — Clinical Data Reasoning (Text-Only) |
| ❌ Text only | deepseek-r1:14b | DeepSeek R1 (14B) — Clinical Reasoning |
| ❌ Text only | openbiollm-70b-chat:Q4_K_M | OpenBioLLM 70B — Comprehensive Biomedical Reasoning |
| ❌ Text only | hf.co/bartowski/baichuan-inc_Baichuan-M2-32B-GGUF:Q6_K | Baichuan M2 32B — Medical Reasoning (HealthBench) |
| ❌ Text only | gpt-oss:120b | GPT-OSS 120B — General Medical Reasoning (OpenAI) |
Request format
Same endpoint as a normal chat, but content becomes an array with a text part and an image_url part (a data URI, or a public image URL) — the standard OpenAI vision format.
curl -s -X POST "BASE_URL/api/chat/completions" \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model": "MODEL_ID",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "What do you see in this image?"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,BASE64_IMAGE_DATA"}}
]
}
]
}'
Verified live: a test image (a blue square, top-left, and a red circle, bottom-right) sent to medgemma-1.5-4b got back — correctly — "I see a blue square and a red circle in the image. The blue square is located in the top left position. The red circle is located in the bottom right position."
Sending an image to a text-only model
You'll get a clean 400, not a silent failure — confirmed live against alibayram/medgemma:4b:
HTTP 400
{"detail":"{\"error\":{\"code\":400,\"message\":\"Multimodal data provided, but model does not support multimodal requests.\",\"type\":\"invalid_request_error\"}}"}
5 Streaming responses
Set "stream": true to get tokens as they're generated, as server-sent events — same shape OpenAI's streaming API uses.
curl -N -X POST "BASE_URL/api/chat/completions" \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model": "MODEL_ID",
"messages": [{"role": "user", "content": "Count from 1 to 5."}],
"stream": true
}'
Each line arrives as data: {json chunk}, ending with data: [DONE] — parse it the same way you'd parse an OpenAI streaming response.
Using an OpenAI client library
Because the API is OpenAI-compatible, you can often just point an existing OpenAI SDK at your instance instead of hand-rolling requests.
from openai import OpenAI
client = OpenAI(
base_url="BASE_URL/api",
api_key="YOUR_TOKEN",
)
resp = client.chat.completions.create(
model="MODEL_ID",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "BASE_URL/api",
apiKey: "YOUR_TOKEN",
});
const resp = await client.chat.completions.create({
model: "MODEL_ID",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(resp.choices[0].message.content);
POST BASE_URL/ollama/api/generate and GET BASE_URL/ollama/api/tags.Notes & troubleshooting
Your instance has to be running
If it's paused or hasn't been opened yet, requests will fail or time out. Start it once from CloudPC.live (or hit any endpoint here — a paused instance won't wake up for API calls the way it does for the web viewer).
401 Unauthorized
Your token expired or was revoked. Sign in again (Authentication) to get a fresh one.
{"detail":"Model not found"}
The most common cause: the model id in your request isn't the one actually loaded on this instance — each instance runs exactly one model at a time (see the note in step 2). Confirmed live: picking a different model in the app's dropdown, or copying a model id from the vision table above, does not change what's running until a new instance is started with that model selected.
Check GET /api/models first — the id you send has to match what's listed there ("loaded": true means it's warm and ready; otherwise the first request may take longer while it loads).
CORS / calling from a browser app
These endpoints are meant to be called from a server, script, or backend you control — not directly from another website's front-end JavaScript with your token embedded in it.