9798183580419 - deep dive into sglang, volume ii: quantization, distributed serving, kernels, and measurement (foundation books) von yang, yin (3 Ergebnisse)

Sprache: Englisch
Verlag: Independently published, 2026
- Softcover
Anbieter: PBShop.store US, Wood Dale, IL, USAPBShop.store US
Verkäufer/-in kontaktierenVerkäufer/-in mit 5 SternenZustand: Neu
EUR 37,38
Versand gratisVersand innerhalb von USAAnzahl: Mehr als 20 verfügbar
PAP. Zustand: New. New Book. Shipped from UK. Established seller since 2000.

Sprache: Englisch
Verlag: Independently published, 2026
- Softcover
Anbieter: PBShop.store UK, Fairford, GLOS, Vereinigtes KönigreichPBShop.store UK
Verkäufer/-in kontaktierenVerkäufer/-in mit 5 SternenZustand: Neu
EUR 30,19
EUR 7,90 VersandVersand von Vereinigtes Königreich nach USAAnzahl: Mehr als 20 verfügbar
PAP. Zustand: New. New Book. Shipped from UK. Established seller since 2000.

Sprache: Englisch
Verlag: Independently Published Jun 2026, 2026
- Softcover
Anbieter: AHA-BUCH GmbH, Einbeck, DeutschlandAHA-BUCH GmbH
Verkäufer/-in kontaktierenVerkäufer/-in mit 5 SternenZustand: Neu
EUR 48,00
EUR 39,01 VersandVersand von Deutschland nach USAAnzahl: 2 verfügbar
Taschenbuch. Zustand: Neu. Neuware - Deep Dive into SGLang, Volume II continues the source-guided explanation of SGLang's inference runtime where the core request path meets deployment pressure.This volume follows the serving contracts introduced in Volume I into quantized weights, LoRA adapters, Mixture-of-Experts routing, dist…ributed placement, collectives, prefill/decode disaggregation, load balancing, custom kernels, CUDA graphs, hardware backends, multimodal transformer serving, diffusion inference, benchmarking, correctness testing, and technical extension work.The book is written for engineers, researchers, and advanced students who want to understand how modern LLM serving systems preserve correctness while changing representation, placement, execution, and measurement.- How rows stay aligned across optimized execution paths- How KV state moves safely between runtime components- How kernels become eligible or ineligible for a request step- How distributed workers coordinate placement and visibility- How benchmark claims become comparable- How extensions can be evaluated without breaking hidden runtime contractsThis is not a command reference or quick-start guide. It is a deep technical reading of the machinery behind high-throughput inference serving. Each chapter explains the algorithmic or systems problem first, then ties SGLang-specific claims to source files, tests, benchmark artifacts, and operational invariants.Volume II assumes familiarity with the core serving loop from Volume I: admission, prefill and decode scheduling, KV storage, forward execution, sampling, streaming, and cleanup. Readers with GPU systems and transformer inference background can also use the opening preliminaries bridge as a compact orientation before entering the specialized chapters.