Branch Divergence - Built-in functions - Shader Learning

Built-in functions

Branch Divergence

Task

Rewrite the shader without using if, in order to avoid warp divergence and preserve parallel execution.

Theory

Let's explore why conditional statements can be dangerous for the GPU and when they're not.


What is a Warp?


A modern GPU consists of many compute units - called Streaming Multiprocessors (SMs) on NVIDIA and Compute Units (CUs) on AMD:



Each unit can manage hundreds or thousands of threads simultaneously. Threads are grouped into blocks or workgroups:



Threads blocks are further divided into warps (NVIDIA) or wavefronts (AMD):



A warp is a group of threads (typically 32 or 64) that execute in lockstep on the GPU. This means all threads in a warp must follow the same instruction path at the same time. Even though each thread may operate on different data, they all follow the same control flow.


To maintain peak performance, each thread in a warp should take the same amount of time to complete its work.


Why branching can be risky


When threads inside a warp encounter an if statement like:

if (condition) {
    result = job_1();
} else {
    result = job_2();
}

and condition evaluates differently across threads, the warp splits:



The GPU must execute both branches. So it runs one branch while disabling threads that don't match, then switches and runs the other. This is called branch divergence and it breaks the warp's parallel efficiency.


What modern GPUs often do


To avoid divergence, modern GPUs may execute both branches anyway, then select the correct result per thread. This is called predicated execution. The above code might be internally transformed into:

vec3 result = mix(job_1(), job_2(), float(condition));

All threads run the same instruction, but both branches are computed.


When it becomes a problem


If both branches contain heavy operations (texture(), loops, expensive math), then executing both can be costly even if only one result is used.


When it is safe


There are exceptions where the GPU knows ahead of time which branch will be taken:


- the condition uses uniforms or constants that are the same across all threads;
- the compiler can statically resolve the condition;
- the warp executes identical logic for all fragments.


In these cases, the GPU can skip one branch entirely - no divergence, no overhead.


Masking vs Branching


Simple if statements and ternary operators like condition ? a : b do not trigger actual branching. Instead, the GPU uses masking to select values without interrupting the execution flow. For example:

float a = (uv.x > 0.5) ? 1.0 : 0.0;

This can be compiled into GPU instructions like:

// compares uv.x with 0.5
cmp_gt_f32 tmp, uv.x, 0.5

// masking
cndmask a, 0.0, 1.0, tmp

There are no jump or branch instructions. cndmask chooses between 0.0 and 1.0 based on tmp, but does not branch, all threads execute the same instruction.


For a deeper dive, Inigo Quilez’s article on GPU conditionals explains how ternary operators are compiled and why they don’t involve branching.


Summary


- if is not inherently bad, but divergence inside a warp breaks parallelism;

- modern GPUs often execute both branches to avoid warp splitting;

- use mix, step, or arithmetic masking for lightweight decisions;

- avoid branching when both paths are computationally expensive;

- uniform-based conditions are safe - the GPU knows what to do.