SGLangを用いたJEV類似決定モデルのスケーリング

SGLangを用いたJEV類似決定モデルのスケーリング

はじめに

顧客が注文の発送状況を問い合わせる際、エージェントはすでに注文IDを把握しており、注文ステータスサービスへの照会、配送ポリシー文書の検索、または顧客への注文ID確認という3つの行動候補に直面する。実行前に選択を行う必要があり、JEV類似決定モデルは長文の説明ではなくカテゴリやスコアを直接返すことで、アプリケーションコードが迅速に行動できるようにする。

決定LLMの本質

決定プロセスは本質的に分類またはスコアリングタスクである。モデルは回答境界の位置において次のトークンのスコアによって判断を下し、説明テキストを生成する必要がない。

2種類のプロンプト方式

ポイントワイズプロンプトは各候補を個別にスコアリングし、セットワイズプロンプトはすべての選択肢をモデルに一度に提示してから判断を下す。

Pointwise prompts isolate each candidate; a setwise prompt places all options before one answer boundary.

専用スコアリングAPIが必要な理由

SGLangの /v1/score APIは特定ラベルトークンのスコアを明示的にリクエストでき、生成的なtop-kによる重要ラベルの見落としを防ぐ。また、単一アイテムスコアリング(SIS)と複数アイテムスコアリング(MIS)の両方をサポートしており、後者は共有クエリの計算を再利用できる。

Grouped bars compare p95 decision latency for 2, 5, 9, and 16 candidates on Qwen3-0.6B, Qwen3-8B, and Qwen3.5-4B.

ベンチマーク:ポイントワイズ決定

Open-Jevデータセットを使用し、単一のNVIDIA H200 GPU上でQwen3-0.6B、Qwen3-8B、Qwen3.5-4Bモデルをテストした。

Panels for Qwen3-0.6B, Qwen3-8B, and Qwen3.5-4B compare Generate, SIS, and MIS p95 latency as the target question rate increases, using a logarithmic latency axis.

結果として、MISは高負荷時に顕著な優位性を示し、レイテンシの増加が緩やかであることが確認された。

The Score API explicitly returns requested labels, while MIS separately enables shared-query execution for independent pointwise candidates.

セットワイズ方式とFused-Choiceの比較

さらに、Fused-ChoiceとSetwise SISをさまざまなQPS条件下でテストした。

Three panels compare Fused-Choice and Setwise SIS p95 latency versus offered QPS on Qwen3-0.6B, Qwen3-8B, and Qwen3.5-4B, with a shared logarithmic latency scale.Grouped bars show Fused-Choice and Setwise SIS p95 latency for 2, 5, 9, and 16 candidates on the three models, using a common zero-based 0 to 60 millisecond scale.

結論:MISは共有クエリのオーバーヘッドを効果的に分散でき、候補数が多いシナリオに適している。