GPU Partitioning vs MIG: What's the Difference

"To share GPUs, don't you just use NVIDIA MIG?" True — but only half true. MIG is powerful, yet it has clear limits that often trip teams up in practice. Here's where MIG falls short and how software-based partitioning addresses it.
First, why partition a GPU?
Assign a full GPU to one small workload and most of it sits idle. The more jobs that use only part of a GPU — dev, test, small inference — the bigger the waste. So splitting one GPU across multiple jobs is key to raising utilization.
The go-to method is NVIDIA's MIG (Multi-Instance GPU). But some things MIG alone can't solve.
What is NVIDIA MIG?
MIG is NVIDIA's technology for dividing one physical GPU into hardware-isolated instances. Each instance gets dedicated SMs (compute cores), dedicated memory, and dedicated cache.
Its strength is hardware isolation — near-zero interference between instances, so one job doesn't affect another's performance. Useful when a multi-tenant setup needs strong isolation.
The catch is the constraints that come with it.
MIG's three constraints
1. Fixed partition sizes. MIG only splits into predefined profiles (1g, 2g, 3g, 7g, etc.). An H100 80GB, for example, splits only into set sizes like 1g.10gb (1/7), 2g.20gb, 3g.40gb. Arbitrary splits like "just 15%" or "0.3 of a GPU" aren't possible. If your workload's size doesn't match a profile, the leftover is wasted.
2. Max seven instances. However finely you split, one GPU tops out at seven instances. Want to run dozens of small models on one GPU? Seven is the ceiling.
3. GPU restrictions. MIG works only on specific data-center GPUs — A100, H100, H200, B200. On other GPUs you can't use it at all.
Aspect | MIG |
|---|---|
Split method | Fixed profiles (predefined sizes) |
Max partitions | 7 |
Arbitrary ratios | No |
Supported GPUs | A100 / H100 / H200 / B200, etc. |
Changing splits | Requires reconfiguration |
How software partitioning differs
The approach that lifts MIG's constraints is software-based partitioning — not bound to hardware profiles, splitting GPU resources more flexibly at the software layer.
TEN's AIPub can split one GPU into up to 100 blocks, at 1% granularity. That means allocating exactly what a workload needs — "15% of a GPU," "37%" — a flexibility impossible with MIG's fixed profiles. It also supports MIG when you want it.
On top of that, AIPub automatically detects idle resources and reallocates them. If partitioned capacity goes unused, it's reclaimed and reassigned where needed — resources flow to actual usage rather than sitting in fixed slices.
Which fits which environment?
The two aren't mutually exclusive. Where hardware-level isolation is essential (e.g., full isolation between tenants at different security levels), MIG's hardware isolation fits. Where you need to finely split many varied-size workloads on one GPU and flexibly reuse idle capacity, software partitioning pushes utilization higher.
The point isn't "can you split a GPU," but "can you split it to fit your workload, without waste."
FAQ
Q. Isn't MIG enough on its own?
If you need strong hardware isolation, MIG fits. But fixed splits, a seven-instance cap, and GPU restrictions limit flexibly dividing many varied workloads.
Q. Can AIPub's 100-block split be used with MIG?
Yes. AIPub supports both software partitioning and MIG, so you can choose or combine per environment.
Q. Isn't software partitioning weaker on isolation?
It's not the physical separation of hardware isolation, but it isolates resources and permissions per team and project to manage interference. Choose the method by the isolation level you need.
Q. Which GPUs can use it?
MIG is limited to certain GPUs (A100, H100, etc.), while software partitioning works across a broader range.
Conclusion: beyond splitting — split to fit
The goal of partitioning isn't just slicing a GPU, but dividing it to fit workloads without waste so expensive GPUs don't idle. MIG offers hardware isolation but comes with fixed splits, a seven-instance cap, and GPU restrictions.
TEN's AIPub flexibly splits GPUs into 100 blocks at 1% granularity, auto-detects and reallocates idle resources, and supports MIG when needed. With per-team/project isolation and real-time monitoring of 50+ metrics in one platform, it answers not just "can you split" but "can you actually raise utilization."
To split GPUs to fit your workloads without waste, explore AIPub.
👉 Learn more about AIPub: https://ten1010.io/
References
NVIDIA, Multi-Instance GPU (MIG) User Guide
NVIDIA, MIG Profiles for A100 / H100 Documentation
TEN, AIPub Product Guide — GPU Partitioning (2026)