OpenJev · Zero-GPU CPU Deploy Kit

⚖️ no GPU · free HF Spaces · exact /v1/systemone contract

Run the OpenJev 27B decision API on pure CPU — the same POST /v1/systemone shape as the hosted openjev-server, without a GPU. Each answer is read off the model's first output token (one forward pass, nothing generated), so a CPU-only 27B is a real, honest API.

🚀 Duplicate the zero-GPU runtime Space or 📦 Get the deploy kit (Docker, compose, docs)

Why this works — no code hacks

openjev-server's vllm backend is really "any OpenAI-compatible server that returns logprobs". When vLLM's logprob_token_ids extension is absent, it falls back to matching top-K logprobs by label text (the documented "old protocol" in openjev_server/backends/vllm.py). llama.cpp speaks exactly that. So this kit runs the official, unmodified server, with llama.cpp (CPU, GGUF) in place of vLLM:

┌─────────────────────────────┐        ┌──────────────────────────────┐
│  openjev-server (official) │        │  llama.cpp llama-server      │
│  :7860  /v1/systemone      │  HTTP  │  :8080  /v1/chat/completions │
│         /v1/version        │───────▶│  logprobs + top_logprobs    │
│         /healthz /readyz   │  JSON  │  (CPU, OpenJev-Q4_K_M.gguf) │
│         /docs /metrics     │        │  16.2 GB, zero CUDA          │
└─────────────────────────────┘        └──────────────────────────────┘

1) Deploy on Hugging Face (free)

  1. Duplicate the runtime Space broadfield/openjev-cpu-run.
  2. In Settings → Hardware: CPU basic (free)  ·  in Settings → Persistent storage: 30 GB (needed for the 16.2 GB model in /data).
  3. Wait for the build + first boot. The image compiles llama.cpp CPU, installs openjev-server, then downloads the Q4 GGUF. ~20–50 min cold; after that restarts load from persistent storage in ~1–2 min.
  4. Your Space is now a live decision API. Try it (don't forget your username in the subdomain):
BASE=https://YOURUSERNAME-openjev-cpu-run.hf.space

curl -s "$BASE/v1/systemone" -H 'Content-Type: application/json' -d '{
  "state": "Customer message: I was charged twice for my order last week and nobody has replied.",
  "questions": {
    "route":   {"type": "choice", "instructions": "Which team should handle this?",
                "criteria": {"billing": null, "shipping": null, "technical": null}},
    "angry":   {"type": "noul",   "instructions": "Is the customer angry?"},
    "urgency": {"type": "score",  "instructions": "How urgent is this?",
                "criteria": ["can wait", "this week", "today", "right now"]}
  }
}'
{
  "answers": {
    "route":   {"type": "choice", "choice": "billing",
                "probabilities": {"billing": 0.9996, "shipping": 0.0002, "technical": 0.0002},
                "confidence": 0.9995},
    "angry":   {"type": "noul", "noul": 0.56},
    "urgency": {"type": "score", "score": 2.04,
                "legend": {"0": "can wait", "1": "this week", "2": "today", "3": "right now"},
                "probabilities": {"0": 0.02, "1": 0.13, "2": 0.64, "3": 0.21},
                "confidence": 0.71}
  },
  "usage": {"input_tokens": 239, "output_tokens": 0}
}

2) Deploy locally / on any VPS

git clone https://huggingface.co/datasets/broadfield/openjev-cpu-deploy-kit
cd openjev-cpu-deploy-kit
docker compose up --build -d
curl -s http://localhost:7860/v1/systemone ...   # same payload as above

plain docker build -t openjev-cpu . + docker run -p 7860:7860 -v openjev-data:/data openjev-cpu also works.

API surface (identical to the hosted server)

endpointmethodwhat
/v1/systemonePOSTany mix of choice / noul / score questions
/v1/prewarmPOSTprefill a long state once, then send its questions
/v1/chat/completionsPOSTtext passthrough (thinking forced off)
/v1/versionGET/POSTmodel, profile, probe, readout code hash
/healthz /readyz /metricsGETliveness, readiness, Prometheus
/docsGETOpenAPI/Swagger UI

Auth: set OPENJEV_TOKEN → all /v1/* require Authorization: Bearer <token>.

3) Ask from code

Python

import httpx

r = httpx.post("http://localhost:7860/v1/systemone", json={
    "state": "Customer message: I was charged twice for my order last week and nobody has replied.",
    "questions": {
        "route": {"type": "choice", "instructions": "Which team should handle this?",
                  "criteria": {"billing": None, "shipping": None, "technical": None}},
        "angry": {"type": "noul", "instructions": "Is the customer angry?"},
    },
}, timeout=60)
answers = r.json()["answers"]
print(answers["route"]["choice"], answers["route"]["probabilities"], answers["angry"]["noul"])

TypeScript

const r = await fetch(`${BASE}/v1/systemone`, {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    state: "Customer message: I was charged twice for my order last week and nobody has replied.",
    questions: {
      route: { type: "choice", instructions: "Which team should handle this?",
               criteria: { billing: null, shipping: null, technical: null } },
      angry: { type: "noul", instructions: "Is the customer angry?" },
    },
  }),
});
const { answers } = await r.json();

Config & tuning

vardefaultwhat it does
MODEL_FILEOpenJev-Q4_K_M.ggufQ4 (16 GB) · Q5 (19 GB) · Q6 (22 GB) · Q8 (28 GB)
MODEL_REPOopenjev/openjev-GGUFwhere the GGUF lives
LLAMA_CTX8192context window
LLAMA_THREADSautopin cores, e.g. 8 / 8
LLAMA_MAX_LOGPROBS64top-K; keep ≥ #options
LLAMA_PARALLEL1concurrent sequences
OPENJEV_TOKEN(empty)bearer auth on /v1/*
OPENJEV_PROFILEopenjevpublished calibration constants

Latency expectations (be honest with yourself)

This is a 27B model on a shared free CPU — not 125 ms. Expect ~2–6 s per single question on a 16-core Space (Q4). But because a decision only ever generates 1 token, batching questions that share one state into a single call makes each extra answer nearly free:

questions / callest. latency @ ~3 SPE
1~4 s
4~4.5 s
10~5.5 s (10 answers, still one pass)

Bench the live Space: node bench.mjs http://…:7860.

Files in this kit

pathroleraw
Dockerfilemulti-stage: llama.cpp CPU + official openjev-server raw
entrypoint.shGGUF → /data, llama-server, then openjev serve raw
docker-compose.ymllocal/VPS one-liner raw
README.mdfull deploy + benchmarks raw
bench.mjsthroughput/latency smoke bench raw

Everything is also mirrored to the dataset repo broadfield/openjev-cpu-deploy-kit.