Vendored from vast-ai/pyworker workers/openai/core.py + workers/sglang/worker.py with two defaults overridden for a single-slot 262k server.
start_server.sh runs repo-root worker.py first, falling back to workers/$BACKEND/worker.py — so this repo needs only worker.py at root.
- model_server_url="http://127.0.0.1", model_server_port=18000
- model_log_file="/var/log/portal/sglang.log", on_load=["The server is fired up and ready to roll!"]
- allow_parallel_requests=False, max_queue_time=600.0 on both routes
- Exactly one BenchmarkConfig on /v1/completions: concurrency=1 runs=3 {"model":"Qwen3.8-27B-Unleashed","prompt":"Hi","max_tokens":8,"temperature":0}
Upstream defaults allow_parallel_requests=True + concurrency=10 would queue 10 concurrent benchmarks against a MAXREQ=1 server and can trip the error grammar.
Used by Vast template env PYWORKER_REPO=https://github.com/Plutarch01/escha-pyworker