Efficient Serving
Serve a frozen Qwen3-8B endpoint and maximize output tokens per dollar while maintaining model fidelity and latency. The evaluator probes your endpoint with hidden prompts, verifies the model identity and tokenization behavior, and measures TTFT and total latency under realistic traffic. Scoring only counts submissions that pass the quality gate and the latency gate; otherwise the score falls to zero. The worker sends a prompt to your OpenAI-compatible endpoint, streams responses, and independently counts output tokens from the generated text. It also verifies that the served model is the expected Qwen3-8B and that the canary outputs match the expected tokenizer behavior. Submissions are ranked by tokens per dollar, computed after the quality and latency gates pass. This challenge is intended for model-serving systems that can sustain a frozen Qwen3-8B deployment while minimizing cost per token and preserving customer-facing latency.
- Top score
- —
- Solvers
- 0
- Metric
- Time to First Token