Monsder guide
Which models fit your GPU?
Select a GPU and the model workload settings, then compare available memory with estimated demand. Monsder's reverse selection helps shortlist candidates; it is not a deployment test and does not verify speed or complete-system compatibility.
Determine available memory
Check the exact GPU variant and memory type. Rated capacity, free memory at startup and memory available to the model are different quantities. Other processes, the runtime and buffers consume some capacity. Shared memory and multiple memory domains need separate placement analysis.
Two 24 GB GPUs do not automatically become one 48 GB device. Distributing a model requires architecture and runtime support; GPU interconnects and layer placement can affect the outcome. Summed capacity alone does not validate a configuration.
Compare under the same conditions
In the hardware-first calculator, set the GPU, mode, precision, context and concurrency. Compare candidates using identical settings. Shorter context or a different weight format can change the estimate, but the new result applies only to those conditions.
A model with an unknown memory profile does not receive a verified fit status. If only weight memory is known, that supports further investigation, not a deployment conclusion. Generation speed cannot be inferred from VRAM capacity or rated peak operations.
What to verify after shortlisting
Retain the exact model, calculation inputs, catalog version and limitations. Check the supported runtime and complete computer configuration. Measure latency, throughput and peak memory on representative requests, including long context and peak concurrency.
If the estimate shows insufficient memory, reconsider the model or explicitly revise the workload conditions. Reduced requirements must stay visible: a candidate sized for a different load does not satisfy the original task by default.
Common questions
Will a model that fits run fast?
That remains unknown until the selected runtime is validated and measured. A memory fit does not establish latency, generation speed, stability or cooling.
Can I check my GPU without registering?
Yes. Open the Monsder calculator and select the hardware-first direction. An account is needed to save a project on the server and request an engineer review.
Sources and scope
Primary documentation explains memory and runtime principles. The Monsder workflow describes the current preliminary calculator. Examples are not GPU benchmarks; final engineering decisions require specialist review.
Apply the method
- How much VRAM does an LLM need?
Estimate LLM memory from weights, quantization, KV cache, context and concurrency. An 8B example and the limits of a preliminary Monsder calculation.
- How to choose an AI server
Choose an AI server from workload requirements: model, load, GPU, memory, power and cooling. What a calculator estimates and what needs engineering validation.