A team must serve a 70B model on a single 80 GB GPU for an internal assistant with modest concurrency. Full FP16 weights will not fit alongside the KV cache for the target context length. They want to keep accuracy loss minimal and are willing to spend additional build time. Which approach best fits these constraints?
Weight-only INT4 quantization shrinks the weight footprint roughly fourfold versus FP16, letting a 70B model fit on one 80 GB GPU while leaving the KV cache and activations in higher precision. Group-wise scales limit accuracy loss, and the extra build and calibration time is acceptable given the stated willingness, making this the best fit.
Why this answer
The binding constraint is fitting a 70B model plus KV cache on one 80 GB GPU while preserving accuracy. Weight-only INT4 quantization cuts weight memory roughly fourfold while keeping activations and cache at higher precision, and group-wise scales with a representative calibration set limit degradation. The team's willingness to accept longer build time matches the calibration and engine-build cost this method requires.
Exam trap
The trap here is reaching for a parallelism or context-shortening workaround when the real constraint is weight footprint on a single device that weight-only low-bit quantization directly addresses.