TL;DR: FP8 vs INT4 quantization has a clear default for most teams: choose FP8 for production LLM workloads where model quality, frequent updates, or mixed-task versatility matter. Reserve INT4 only when hard memory ceilings or cost constraints on stable, well-characterized workloads make the extra engineering overhead worthwhile. As of 2026, FP8 on Hopper-generation hardware has narrowed the gap with BF16 enough to treat it as the new production baseline.
Key takeaways
- INT4 accuracy loss is task-specific: Reasoning and code generation take measurable hits; retrieval tasks hold up better.
- INT4 carries a real engineering tax: Every model update requires a new GPTQ or AWQ calibration run plus regression testing.
- FP8 should be your default: Unless hard memory limits force you lower, FP8's accuracy and zero calibration overhead make it the lower-risk production choice.
- Mixed-format is a practical middle ground: FP8 weights plus INT4 KV cache recovers memory headroom without compressing the component that most directly affects output quality.
- Benchmark on your workload: Published accuracy numbers rarely match what a specific production task sees.
Does FP8 work on H100?
FP8 executes natively on Hopper-generation GPUs, including the H100, making it a first-class serving format rather than a compromise. The quantization community broadly treats FP8 as the safest option, saving significant memory with almost no visible accuracy loss. On Hopper hardware, the operational gap between BF16 and FP8 has narrowed to the point where FP8 warrants no special justification, treat it as the new baseline and require justification to go lower.
How does FP8 hardware support on Hopper change the quantization calculus?
FP8 native support on Hopper-generation GPUs removes the primary technical objection to adopting it as a production default. One hard constraint remains: FP8 is not supported on Ampere-generation GPUs such as the RTX 3090. If your cluster is Ampere, INT4 or INT8 are your practical ceiling. On Hopper hardware, there is no hardware reason to drop below FP8. Require justification to go lower rather than the reverse.

Where does INT4 accuracy degradation actually hurt, and where can you tolerate it?
INT4 degrades reasoning and code generation measurably but performs acceptably on knowledge retrieval tasks. Your workload type, not a generic benchmark, determines whether the accuracy loss matters for your deployment.
INT4 maps weights to just 16 integer values, compressing the dynamic range that multi-step reasoning depends on most. INT4 degrades code generation more than knowledge tasks, that asymmetry is your first filter.
Table 1: FP8 vs INT4 Workload Accuracy Risk Framework
| Workload type | INT4 accuracy risk | FP8 accuracy risk | Recommended format |
|---|---|---|---|
| Multi-step reasoning | High | Negligible | FP8 |
| Code generation | High | Negligible | FP8 |
| Knowledge retrieval / RAG | Low-medium | Negligible | FP8 or INT4 |
| Classification / routing | Low | Negligible | INT4 viable |
| Long-context summarization | Medium | Negligible | FP8 preferred |
A coding copilot and a document-search assistant running on identical hardware reach opposite conclusions. If your primary workload sits in the first two rows of Table 1, the accuracy evidence makes INT4 very difficult to justify before any further benchmark is run.
What is the true total cost of INT4 deployment when you factor in calibration and model refresh cycles?
The engineering overhead of INT4 is real and recurring: every model update requires a new calibration run, regression testing, and sign-off before the update can ship. Whether that labor cost outweighs the energy and memory savings depends on your specific model churn rate and available headcount, not on any universal rule.
The energy efficiency case for INT4 is real. MXINT8 and NVINT4 reduce energy by 37-38% compared with MXFP8 and NVFP4, and Intel's INT4 variant offers the best overall trade-off among INT4 variants on Qwen3.6 27B. As a practical rule of thumb used in this guide: if your team ships model updates frequently, the calibration cost compounds and should be treated as a sustained staffing commitment, not a one-time setup expense.

When does mixed-format quantization make sense as a production strategy?
Pairing FP8 weights with an INT4 KV cache is the most defensible middle path when you need memory headroom beyond what FP8 alone provides but cannot accept the full accuracy and calibration cost of pure INT4 weights.
This guide calls this strategy the weight-cache precision split: apply FP8 where the model is most sensitive to precision loss and INT4 where it is least. Weights encode the model's learned parameters. The KV cache holds attention state for the current context window. Based on the author's reading of available literature, the KV cache is the lower-risk target for aggressive compression, though this classification should be treated as author synthesis rather than a universally measured result.
The real win from quantization is concurrency throughput, not raw per-request speed. FP8 weights preserve output quality while INT4 KV cache recovers memory that lets you serve more concurrent requests. Mixed-format still requires workload-specific validation before deployment.
Frequently asked questions
What is the main difference between FP8 and INT4 quantization? FP8 uses floating-point representation at 8 bits, preserving a wider dynamic range. INT4 maps weights to just 16 integer values, maximizing compression at the cost of precision. FP8 retains more accuracy than INT8 in certain scenarios and requires Hopper-generation hardware, while INT4 is well-supported across all hardware including older Ampere-class GPUs.
When does INT4 accuracy loss become unacceptable for production reasoning workloads? When the workload requires multi-step reasoning, code generation, or long-context coherence. INT4 degrades code generation more than knowledge tasks, and calibration tuning does not recover that loss. If those workloads are your primary use case, the accuracy evidence weighs heavily against INT4 before evaluation begins.
How should teams benchmark quantization quality loss before committing to INT4? Run task-specific evaluations on your actual production prompts, not generic benchmarks like MMLU. Score outputs from both formats against your own quality rubric, and treat any degradation above your SLA threshold as a hard block. The benchmark that matters is the one on your data.
Can FP8 replace BF16 entirely for production inference on H100? For most inference workloads, the evidence suggests yes. FP8 loses no measurable accuracy on Hopper hardware and is described as "the safest option with almost no visible accuracy loss." The framework used in this guide treats FP8 as a direct inference replacement for BF16 on Hopper, with a meaningful reduction in memory footprint.
Conclusion
For teams that need memory headroom beyond FP8 alone, the weight-cache precision split, FP8 weights paired with INT4 KV cache, is the named middle path this guide recommends before committing to full INT4. Pick your format based on model churn rate and concurrency target, not the memory savings headline. Run your production prompts through both formats before committing. The benchmark that matters is the one on your data.
Learn from me

Forward Deployed Engineering Bootcamp for Full-Stack Developers, my Maven cohort. Build and ship complete AI products end to end, from React and Node.js frontends to deployed models with caching and observability. Join the next cohort →
Hire us
Traversaal.ai. We're a team of forward deployed engineers solving the toughest AI problems for Fortune 100 companies: document intelligence, agentic data platforms, and real-time web intelligence, deployed in production. Work with our team to deploy your next agentic ecosystem. Talk to Traversaal.ai →
Join us
Want to solve these problems with us? We're always looking for forward deployed engineers who want to ship production AI. jobs@traversaal.ai
