Isbn: 9798189579370 - fast and frugal llm apps: systematic cost and latency optimization with caching, routing, and model right-sizing (applied llm engineering series, band 10) (3 Ergebnisse)

ISBN: 
Mit der Detailsuche verfeinern

Optimieren Sie Ihre Suche

  • Bücher (3)

  • Neu (3)

bis

Benutzerdefinierte Preisspanne (EUR)

bis

  • Sprache: Englisch

    Verlag: Amazon Digital Services LLC - Kdp, 2026

    9798189579370

    Serie: Buch 10 von 12 - Applied LLM Engineering Series

    • Softcover

    Anbieter: PBShop.store US, Wood Dale, IL, USAPBShop.store US

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 14,24

     Versand gratis 
    Versand innerhalb von USA

    Anzahl: Mehr als 20 verfügbar

    PAP. Zustand: New. New Book. Shipped from UK. Established seller since 2000.

  • Sprache: Englisch

    Verlag: Amazon Digital Services LLC - Kdp, 2026

    9798189579370

    Serie: Buch 10 von 12 - Applied LLM Engineering Series

    • Softcover

    Anbieter: PBShop.store UK, Fairford, GLOS, Vereinigtes KönigreichPBShop.store UK

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 13,10

    EUR 3,87 Versand 
    Versand von Vereinigtes Königreich nach USA

    Anzahl: Mehr als 20 verfügbar

    PAP. Zustand: New. New Book. Shipped from UK. Established seller since 2000.

  • Sprache: Englisch

    Verlag: Amazon Digital Services LLC - Kdp Jul 2026, 2026

    9798189579370

    Serie: Buch 10 von 12 - Applied LLM Engineering Series

    • Softcover

    Anbieter: AHA-BUCH GmbH, Einbeck, DeutschlandAHA-BUCH GmbH

    Verkäufer/-in mit 5 Sternen
    Verkäufer/-in kontaktieren

    Zustand: Neu

    EUR 13,16

    EUR 35,00 Versand 
    Versand von Deutschland nach USA

    Anzahl: 2 verfügbar

    Taschenbuch. Zustand: Neu. Neuware - Stop Letting LLM Inference Bills Drain Your Product's MarginsIn the modern AI landscape, model inference spend has skyrocketed to become the dominant line item for software teams. Many organizations attempt to optimize naively, resulting in a fragile, hard-to-maintain stack of half-finished techniques. Fast and Frugal LLM Apps replaces this chaotic pattern with a single, repeatable engineering framework designed to systematically slash costs and latency without sacrificing output quality.Written specifically for AI engineers, backend developers, engineering managers, and technical founders, this book provides a production-ready playbook to scale your AI systems sustainably. You will learn to treat cost as a first-class engineering metric alongside latency and accuracy.What you will master inside: - The Four-Phase Loop: Run a highly effective six-week optimization cadence with clear ownership.- Smart Caching Architecture: Implement prompt and semantic caching to slash input costs by 70-90%.- Dynamic Routing: Automatically classify requests to send them to the cheapest sufficient model tier.- Open-Source and Quantization: Deploy highly efficient open-weight models in hybrid architectures.- Cost-Aware Evaluations: Establish automated regression gates so optimizations never silently destroy quality.- Batch and Latency Engineering: Transition non-interactive tasks to batch APIs and implement smart streaming.Stop guessing your monthly API invoice. Build a structured cost-engineering practice with named owners, finance partnerships, and solid quarterly reviews. Scroll up, click 'Buy Now,' and transform your model invoice from an unpredictable crisis into a predictable engineering roadmap today. …