A GPU Too Busy to Serve Its Own Owner
When one AI job fills an entire GPU, a small urgent job can be stuck waiting behind it. That is the situation NVIDIA's CUDA green contexts are designed to fix, and as of CUDA 13.1 the tool is available to a much wider range of software. The mechanism lets one application draw a hard line through the GPU's execution resources, reserving a fixed slice of the chip for critical work so that a flood of background computation cannot push it aside.
The primary event for this analysis is an NVIDIA Technical Blog post published October 6, 2026, titled "Control How Your GPU Shares Work with Green Contexts," written by Kyrylo Perelygin, Shelton Dsouza and Myrto Papadopoulou. NVIDIA, the company that designs the chips and writes the CUDA software, is both the vendor and the source of the performance numbers discussed here. Every benchmark figure below is vendor-reported and has not been independently verified.
A note on that source page: the blog carries an AI-generated summary banner, produced by NVIDIA's Nemotron system, with an explicit warning that "AI-generated content may summarize information incompletely." This article does not rely on that banner. All claims below are drawn from the human-written body of the post and from NVIDIA's official CUDA documentation pages.
Green contexts are not brand new. They have existed in the CUDA Driver API since CUDA 12.4, which is a lower-level interface that relatively few applications touch directly. What changed with CUDA 13.1 is access through the CUDA Runtime API, the standard library that ordinary CUDA programs use. NVIDIA's documentation states plainly: "Starting from CUDA 13.1, contexts are exposed in the CUDA runtime via the execution context (EC) abstraction." In practical terms, mainstream applications can now partition a GPU with a small amount of host-side code and no kernel changes at all.
Why Priority Alone Runs Out of Road
The problem green contexts solve has a simple shape. A modern GPU application is often not one program but several running at once inside the same process: a model inference stage, a data preprocessing stage, a communication routine, or a small latency-critical operator sitting next to a large throughput-oriented worker. In CUDA's traditional model, the developer can hint at importance using stream priority, but hints have a hard limit.
NVIDIA's blog post explains that limit precisely: "Stream priorities can't preempt a block that's already executing on an SM, so it needs to wait for some of the blocks to drain." An SM is a streaming multiprocessor, the chip's parallel workhorse. If a bulk job's blocks sit on every SM, a priority bump for the urgent kernel does not evict them. The urgent kernel waits for the bulk job to finish blocks, no matter how high its priority is.
Green contexts attack the problem from the other side. Instead of trying to jump the queue, the application creates lightweight contexts, each provisioned with a specific subset of the GPU's streaming multiprocessors and, separately, its workqueue resources. Work submitted through a green context can only run on the SMs it was provisioned with, regardless of the kernel's launch configuration. As NVIDIA's CUDA Programming Guide puts it, when a latency-sensitive kernel is launched it "is guaranteed that there will be available SMs for it to start executing immediately, barring any other resource constraints."
The urgent job never waits because its slice of the chip is always free. The cost is equally direct: the bulk job gets fewer SMs. That is the entire tradeoff, and it is a deliberate one, set by the developer when the partition is created.
The Vendor-Reported Numbers
NVIDIA's benchmark, reported in the October 6 blog post, uses an NVIDIA Blackwell GPU with 148 SMs. The test runs a tiny critical integer kernel while twenty back-to-back launches of a large bulk kernel saturate the device, and measures the critical kernel's wall-clock latency in three modes: a green-context partition with 8 SMs dedicated to the critical kernel and 140 to the bulk job; a default context with a high-priority critical stream and no partition; and a default context with equal-priority streams and no partition.
All three numbers are vendor-reported, from NVIDIA's own blog post, and have not been independently verified. The reported results were 0.007 ms for the green-context partition, 0.140 ms for the high-priority stream without a partition, and 3.727 ms for equal-priority streams. The post's framing is that stream priority alone was roughly 27 times faster than equal-priority streams, and that a partition added roughly another 20 times improvement on top of priority, because the urgent kernel skipped the wait for bulk work to drain entirely. The cost side of the ledger is that the bulk kernel ran with 140 SMs instead of all 148.
Two caveats deserve emphasis for any reader considering the technology. First, the workload is a synthetic one designed by the vendor to show the mechanism at its clearest, with a critical kernel small enough to fit in 8 SMs. Second, the partition granularity is not arbitrary: NVIDIA's documentation specifies that on compute architectures 7.x and 8.x, SM counts must be multiples of two, and on 9.0 and later, multiples of eight by default, with the exact alignment queryable from the device. Applications also need enough headroom, since a partition only makes sense on devices where the critical job can plausibly live in its reserved slice.
What Actually Changed in 13.1
The programming change is small, and that is the point of the 13.1 exposure. In the traditional model, an application calls cudaSetDevice, creates streams, and the stream's execution target is inferred from thread-local device state. With green contexts, the application instead calls cudaGreenCtxCreate with a resource descriptor describing the SM and workqueue partition it wants, then creates the stream from that context using cudaExecutionCtxStreamCreate. Work launched on the stream lands on the partitioned resources automatically.
The runtime API reference describes cudaExecutionContext_t as an abstraction that can represent either the primary context, which runtime users have always interacted with implicitly, or a green context. The programming guide makes the migration story explicit for existing code: "Using green contexts does not require any GPU code (kernel) changes, just small host-side changes." Applications that do not adopt the feature behave exactly as before; NVIDIA calls the model "opt-in and additive."
There is also an explicit no-op path for developers who want the new handle-based style without partitioning. The runtime provides cudaDeviceGetExecutionCtx, which returns the execution context corresponding to a device's primary context, so code can uniformly target either the full device or a partition through the same API shape.
Where Green Contexts Fit Next to MIG and MPS
How does this relate to NVIDIA's existing partitioning mechanisms? The CUDA Programming Guide draws the comparison directly. Multi-Instance GPU, or MIG, statically slices a supported GPU into several smaller GPUs before an application ever starts, which suits cloud providers running different clients, but it cannot fix intra-process contention, because one application on a single MIG instance can still occupy all the SMs of that instance. The Multi-Process Service, MPS, shares a GPU across multiple processes and, starting with CUDA 13.1, can also statically partition SMs per process.
The distinguishing feature of green contexts is scope: they partition resources within a single process, at creation time, with lightweight contexts that NVIDIA says are much cheaper to create than MPS contexts because many underlying structures are shared with the primary context. The guide's summary of the mechanism boundary is worth keeping in mind: MIG and MPS divide between applications; green contexts divide between components of one application, though the two approaches can be combined, with green contexts partitioning the SMs available inside a MIG instance.
A canonical motivating workload named in the blog is sensor processing, specifically latency-sensitive operators in platforms like NVIDIA's Holoscan that need to start as soon as possible. The blog was also cited for a second common case, overlapping a communication kernel with a GEMM, the matrix-multiply engine of distributed training, so both make progress instead of contending.
Access and the Right to Be Responsive
There is an access dimension here that goes beyond chip mechanics. On shared and edge hardware, whether a cloud GPU, a robot's onboard computer, or an industrial sensor pipeline, the applications that most need guaranteed response are often the smallest ones. Under best-effort scheduling, a small perception operator can be starved by whoever else happens to be resident. Explicit partitioning means a latency-critical component can hold a guaranteed slice of shared silicon instead of hoping a priority flag is honored in practice.
That said, the guarantee is intra-process and intra-device. Green contexts do not, by themselves, arbitrate between different tenants of a cloud GPU; that remains the territory of MIG and MPS. What they do is give individual developers a concrete, documented mechanism to defend their own critical work against their own bulk work, which is a genuine improvement in who can control their application's fate on a shared chip.
What Is Fact, What Is Prediction
The observed facts are straightforward: the APIs exist in CUDA 13.1's Runtime and 12.4-era Driver interfaces, the code paths and documentation are published, and NVIDIA reports a benchmark on a 148-SM Blackwell GPU showing a two-orders-of-magnitude latency gap between a partitioned critical kernel and an equal-priority unpartitioned one. The vendor's own figures should be treated as accurate descriptions of its test, not as universal outcomes; results will depend on workload shape, SM alignment constraints, and other resource dependencies the documentation explicitly flags as unresolved.
The predictions are less certain. Whether distributed-training frameworks, sensor platforms, or cloud middleware adopt green contexts widely is not yet observable, and adoption will hinge on how much throughput teams are willing to give up to buy latency certainty. The technology itself is well documented and low-friction to adopt; whether the tradeoff is worth it in production is a question only real workloads, not vendor benchmarks, can answer.
