9798182198233 - deep dive into sglang, volume i: runtime, scheduling, memory, and decoding (foundation books) von yang, yin (3 Ergebnisse)

Sprache: Englisch
Verlag: Independently published, 2026
- Softcover
Anbieter: PBShop.store US, Wood Dale, IL, USAPBShop.store US
Verkäufer/-in kontaktierenVerkäufer/-in mit 5 SternenZustand: Neu
EUR 37,56
Versand gratisVersand innerhalb von USAAnzahl: Mehr als 20 verfügbar
PAP. Zustand: New. New Book. Shipped from UK. Established seller since 2000.

Sprache: Englisch
Verlag: Independently published, 2026
- Softcover
Anbieter: PBShop.store UK, Fairford, GLOS, Vereinigtes KönigreichPBShop.store UK
Verkäufer/-in kontaktierenVerkäufer/-in mit 5 SternenZustand: Neu
EUR 30,19
EUR 7,90 VersandVersand von Vereinigtes Königreich nach USAAnzahl: Mehr als 20 verfügbar
PAP. Zustand: New. New Book. Shipped from UK. Established seller since 2000.

Sprache: Englisch
Verlag: Independently Published Jun 2026, 2026
- Softcover
Anbieter: AHA-BUCH GmbH, Einbeck, DeutschlandAHA-BUCH GmbH
Verkäufer/-in kontaktierenVerkäufer/-in mit 5 SternenZustand: Neu
EUR 48,00
EUR 39,23 VersandVersand von Deutschland nach USAAnzahl: 2 verfügbar
Taschenbuch. Zustand: Neu. Neuware - Deep Dive into SGLang, Volume I explains SGLang's runtime request path as a set of algorithms, data structures, and systems tradeoffs. This volume focuses on the core inference-serving loop: tokenization, admission, prefill and decode scheduling, prefix lookup, KV-cache allocation, forward ex…ecution, sampling, streaming, and cleanup. Instead of treating SGLang as a catalog of commands, it builds a working model of the runtime state that makes modern LLM serving possible. Volume I covers: - The inference serving problem and request lifecycle- Transformer inference cost models- SGLang's runtime as a distributed state machine- Continuous batching and chunked prefill- KV-cache memory management- RadixAttention and prefix reuse- Hierarchical caching- Attention backends and forward batches- Model architecture registration- Sampling and logits processing- Structured outputs, reasoning parsers, and tool-call parsing- Speculative decoding- Diffusion language models and blockwise decoding>This is Volume I of Deep Dive into SGLang. Volume II continues into quantization, distributed serving, kernels, hardware backends, benchmarking, correctness, multimodal serving, diffusion inference, and technical extension. Independent explanatory guide. Not affiliated with or endorsed by the SGLang project.