Skip to content

Response Time ​

Response time is the total duration an item spends in the system from the instant it arrives until it completes service — including waiting in queue and active processing. It is the primary quality indicator experienced by end users.


Theoretical Foundation: Little's Law ​

You do not need to stopwatch individual tokens moving through the model. For any stable queueing system in statistical equilibrium, Little's Law holds:

$$N = \lambda \cdot W \implies W = \frac{N}{\lambda}$$

In practical terms:

Response Time = (Total items present in the system) ÷ (System throughput)

If an average of 10 customers are present in a branch and the branch finishes 2 customers per minute, the average time each customer spends in the branch is $10 / 2 = 5$ minutes.


Step 1: Total Items in System ​

Sum the time averages of all places through which items pass between arrival and departure:

E{#Queue} + E{#InService}

Include all intermediate stages: waiting, active execution, and network transit. Omitting any place leads to an underestimation of response time.


Step 2: System Throughput ​

There are two approaches to compute the divisor:

Approach 1: Divide by Measured Throughput ​

(E{#Queue} + E{#InService}) / measured_throughput

Approach 2: Multiply by Arrival Delay ​

In a stable system without losses, throughput equals the arrival rate ($\lambda = 1 / \text{ARRIVAL_DELAY}$). Therefore, dividing by $\lambda$ is algebraically identical to multiplying by ARRIVAL_DELAY:

(E{#Queue} + E{#InService}) * ARRIVAL_DELAY

Why Approach 2 is Far More Robust ​

In stationary analysis, both formulations converge to virtually identical numbers:

FormulationStationary Result
Approach 1 (divide by measured throughput)10.04 min
Approach 2 (multiply by arrival delay)10.03 min

However, in transient analysis, Approach 1 breaks down completely:

Time Instant ($t$)Approach 1 (divide by measured throughput)Approach 2 (multiply by arrival delay)
$t = 25$0.007.96 min
$t = 50$0.008.80 min
$t = 100$0.0010.92 min
$t = 200$0.0011.92 min

Cause of Transient Failure

Transient instantaneous throughput measured from short-lived departure places drops near zero in most replications, triggering zero-division guards. Approach 2 depends only on a known constant parameter (ARRIVAL_DELAY), producing a smooth, mathematically sound curve across all regimes.


Systems with Dropping or Capacity Limits ​

When a system has finite capacity and drops arriving items when full, only a fraction of arrivals enter the system. The effective arrival rate is scaled by $(1 - P{\text{full}})$:

(E{#Queue} + E{#InService}) * ARRIVAL_DELAY / (1 - P{#Slots = 0})

This expression correctly evaluates the average response time for items that were admitted.

Always interpret alongside blocking

Constraining queue capacity shortens the response time of admitted users by shedding load. Always report response time alongside the blocking probability.

Next steps ​

Learn how to evaluate resource utilization and bottlenecks in Utilization.