Skip to content

Utilization ​

Utilization ($U$) measures the percentage of time a resource remains busy servicing work. It is the definitive metric for locating bottlenecks: across a chain of workflow stages, the resource with the highest utilization governs the overall capacity limit of the entire system.


Formulating the Metric ​

1. Direct Form (Active Tokens) ​

Divide the average number of items in service by the total number of resource units and multiply by 100:

(E{#InService} / NUM_WORKERS) * 100

2. Complementary Form (Idle Tokens) ​

If your model maintains an explicit place for idle resources, compute the complement of idleness:

(1 - E{#IdleWorkers} / NUM_WORKERS) * 100

Both formulas yield mathematically identical values.


Interpreting Utilization Ranges ​

Utilization RangeOperational Health
Below 50%Resource has comfortable headroom. Queues rarely form.
50% to 70%Balanced target range recommended for steady production systems.
70% to 85%Caution zone: queues and wait times begin growing exponentially.
Above 85%Critical zone: minor random arrival fluctuations produce severe delay spikes.

Why sizing for 95% utilization is dangerous

In queueing theory, response time does not degrade linearly: it explodes asymptotically as utilization approaches 100%:

Worker UtilizationMean Response Time
50%10.0 min
70%23.3 min
80%40.0 min
90%89.5 min

Designing an architecture to run at 90% or 95% load to "maximize efficiency" guarantees unacceptably degraded user latency.


Identifying System Bottlenecks ​

Declare a utilization metric for every component in your model (CPU, database connection pool, network interfaces, worker threads). After running the simulation:

  1. The component with the highest utilization rate is the active bottleneck.
  2. Increasing the capacity of any other component will yield zero overall throughput gain, merely leaving non-bottleneck resources more idle.
  3. Once the primary bottleneck is expanded, rerun the analysis: the bottleneck frequently shifts to the next most utilized stage in the pipeline.

If utilization does not rise with increased load

If incoming load increases but a downstream component's utilization remains flat, an upstream stage is starving it (upstream blocking or accidental single-server configuration).

Next steps ​

Learn how to evaluate rejection rates in Probability / Blocking.