Run the OpenJev 27B
decision API on pure CPU — the same
POST /v1/systemone shape as the hosted
openjev-server,
without a GPU. Each answer is read off the model's first output token
(one forward pass, nothing generated), so a CPU-only 27B is a real, honest API.
openjev-server's vllm backend is really
"any OpenAI-compatible server that returns logprobs". When vLLM's
logprob_token_ids extension is absent, it falls back to
matching top-K logprobs by label text (the documented "old protocol"
in openjev_server/backends/vllm.py). llama.cpp speaks exactly
that. So this kit runs the official, unmodified server, with
llama.cpp (CPU, GGUF) in place of vLLM:
┌─────────────────────────────┐ ┌──────────────────────────────┐
│ openjev-server (official) │ │ llama.cpp llama-server │
│ :7860 /v1/systemone │ HTTP │ :8080 /v1/chat/completions │
│ /v1/version │───────▶│ logprobs + top_logprobs │
│ /healthz /readyz │ JSON │ (CPU, OpenJev-Q4_K_M.gguf) │
│ /docs /metrics │ │ 16.2 GB, zero CUDA │
└─────────────────────────────┘ └──────────────────────────────┘
/data).BASE=https://YOURUSERNAME-openjev-cpu-run.hf.space
curl -s "$BASE/v1/systemone" -H 'Content-Type: application/json' -d '{
"state": "Customer message: I was charged twice for my order last week and nobody has replied.",
"questions": {
"route": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": null, "shipping": null, "technical": null}},
"angry": {"type": "noul", "instructions": "Is the customer angry?"},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["can wait", "this week", "today", "right now"]}
}
}'
{
"answers": {
"route": {"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.9996, "shipping": 0.0002, "technical": 0.0002},
"confidence": 0.9995},
"angry": {"type": "noul", "noul": 0.56},
"urgency": {"type": "score", "score": 2.04,
"legend": {"0": "can wait", "1": "this week", "2": "today", "3": "right now"},
"probabilities": {"0": 0.02, "1": 0.13, "2": 0.64, "3": 0.21},
"confidence": 0.71}
},
"usage": {"input_tokens": 239, "output_tokens": 0}
}
git clone https://huggingface.co/datasets/broadfield/openjev-cpu-deploy-kit
cd openjev-cpu-deploy-kit
docker compose up --build -d
curl -s http://localhost:7860/v1/systemone ... # same payload as above
plain docker build -t openjev-cpu . +
docker run -p 7860:7860 -v openjev-data:/data openjev-cpu also works.
| endpoint | method | what |
|---|---|---|
/v1/systemone | POST | any mix of choice / noul / score questions |
/v1/prewarm | POST | prefill a long state once, then send its questions |
/v1/chat/completions | POST | text passthrough (thinking forced off) |
/v1/version | GET/POST | model, profile, probe, readout code hash |
/healthz /readyz /metrics | GET | liveness, readiness, Prometheus |
/docs | GET | OpenAPI/Swagger UI |
Auth: set OPENJEV_TOKEN → all /v1/* require
Authorization: Bearer <token>.
import httpx
r = httpx.post("http://localhost:7860/v1/systemone", json={
"state": "Customer message: I was charged twice for my order last week and nobody has replied.",
"questions": {
"route": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": None, "shipping": None, "technical": None}},
"angry": {"type": "noul", "instructions": "Is the customer angry?"},
},
}, timeout=60)
answers = r.json()["answers"]
print(answers["route"]["choice"], answers["route"]["probabilities"], answers["angry"]["noul"])
const r = await fetch(`${BASE}/v1/systemone`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
state: "Customer message: I was charged twice for my order last week and nobody has replied.",
questions: {
route: { type: "choice", instructions: "Which team should handle this?",
criteria: { billing: null, shipping: null, technical: null } },
angry: { type: "noul", instructions: "Is the customer angry?" },
},
}),
});
const { answers } = await r.json();
| var | default | what it does |
|---|---|---|
MODEL_FILE | OpenJev-Q4_K_M.gguf | Q4 (16 GB) · Q5 (19 GB) · Q6 (22 GB) · Q8 (28 GB) |
MODEL_REPO | openjev/openjev-GGUF | where the GGUF lives |
LLAMA_CTX | 8192 | context window |
LLAMA_THREADS | auto | pin cores, e.g. 8 / 8 |
LLAMA_MAX_LOGPROBS | 64 | top-K; keep ≥ #options |
LLAMA_PARALLEL | 1 | concurrent sequences |
OPENJEV_TOKEN | (empty) | bearer auth on /v1/* |
OPENJEV_PROFILE | openjev | published calibration constants |
This is a 27B model on a shared free CPU — not 125 ms. Expect ~2–6 s per single question on a 16-core Space (Q4). But because a decision only ever generates 1 token, batching questions that share one state into a single call makes each extra answer nearly free:
| questions / call | est. latency @ ~3 SPE |
|---|---|
| 1 | ~4 s |
| 4 | ~4.5 s |
| 10 | ~5.5 s (10 answers, still one pass) |
Bench the live Space: node bench.mjs http://…:7860.
| path | role | raw |
|---|---|---|
Dockerfile | multi-stage: llama.cpp CPU + official openjev-server | raw |
entrypoint.sh | GGUF → /data, llama-server, then openjev serve |
raw |
docker-compose.yml | local/VPS one-liner | raw |
README.md | full deploy + benchmarks | raw |
bench.mjs | throughput/latency smoke bench | raw |
Everything is also mirrored to the dataset repo broadfield/openjev-cpu-deploy-kit.