Hasta 40% off y Envío a todo USA y PR por solo $2.99  Ver más

Enviar a
FL
0
es
  • argentina
  • chile
  • colombia
  • españa
  • méxico
  • perú
  • estados unidos
  • internacional

Selecciona tu país

América

Europa

Resto del mundo

Idioma
esEspañolActual
enEnglish
portada Quantized Model Deployment. INT8 and FP16 Compression for Mobile Acceleration (en Inglés)
Formato
Libro Físico
Año
2026
Idioma
Inglés
N° páginas
234
Encuadernación
Tapa Blanda
Dimensiones
9.61x6.69x0.47 in
ISBN13
9798196245466

Quantized Model Deployment. INT8 and FP16 Compression for Mobile Acceleration (en Inglés)

Clara Whiskers (Autor) · Independently published · Tapa Blanda

Quantized Model Deployment. INT8 and FP16 Compression for Mobile Acceleration (en Inglés) - Clara Whiskers

Libro Nuevo Origen: Estados Unidos
Envío: 7 a 9 días háb.
$ 19.99$ 17.64
-12%
Libro Nuevo

Quedan más de 100 unidades

$ 17.64
Llega entre el 25 Sep y el 01 Oct a FL. Seleccionar ubicación

Reseña del libro "Quantized Model Deployment. INT8 and FP16 Compression for Mobile Acceleration (en Inglés)"

What if the only thing standing between your neural network and real-time mobile performance is the precision you refuse to give up?
Your model ran flawlessly in PyTorch-400MB of FP32 weights, a 350-watt GPU, and all the thermal headroom in the world. Then you deployed it to a phone. It stuttered. It heated up. The OS killed it before it produced a single inference. The market no longer asks whether AI can run on mobile. It asks why your AI is slower and less accurate than the cloud version. The answer is not your architecture. It is your precision.
This book is the field manual for engineers who refuse to accept the old compromise of smaller models and weaker accuracy. Inside, you will learn:
• Why INT8 and FP16 are not arbitrary format choices, but hardware-mandated keys to dedicated acceleration paths on Snapdragon, Apple Neural Engine, and MediaTek APU • How naïve post-training quantization can crater accuracy by double-digit percentages-and the calibration, range estimation, and outlier handling techniques that prevent it • The exact deployment architecture for TensorFlow Lite, Core ML, ONNX Runtime Mobile, and NNAPI, including operator fusion and numerical equivalence testing • Why quantization is the only optimization that simultaneously improves latency, accuracy, and power consumption-and how to combine it with pruning and knowledge distillation for wearables and IoT
Stop accepting the compromise between speed and accuracy. Build models that run cooler, faster, and sharper on the devices already in your users' pockets. The precision you can no longer afford is the precision you can finally reclaim.

Opiniones del libro

Preguntas frecuentes sobre el libro

Todos los libros de nuestro catálogo son Originales.
El libro está escrito en Inglés.
La encuadernación de esta edición es Tapa Blanda.

Preguntas y respuestas sobre el libro

¿Tienes una pregunta sobre el libro? Inicia sesión para poder agregar tu propia pregunta.

Opiniones sobre Buscalibre

Ver más opiniones de clientes