Theme
Response Time
Response time is the total duration an item spends in the system from the instant it arrives until it completes service — including waiting in queue and active processing. It is the primary quality indicator experienced by end users.
Theoretical Foundation: Little's Law
You do not need to stopwatch individual tokens moving through the model. For any stable queueing system in statistical equilibrium, Little's Law holds:
$$N = \lambda \cdot W \implies W = \frac{N}{\lambda}$$
In practical terms:
Response Time = (Total items present in the system) ÷ (System throughput)
If an average of 10 customers are present in a branch and the branch finishes 2 customers per minute, the average time each customer spends in the branch is $10 / 2 = 5$ minutes.
Step 1: Total Items in System
Sum the time averages of all places through which items pass between arrival and departure:
E{#Queue} + E{#InService}Include all intermediate stages: waiting, active execution, and network transit. Omitting any place leads to an underestimation of response time.
Step 2: System Throughput
There are two approaches to compute the divisor:
Approach 1: Divide by Measured Throughput
(E{#Queue} + E{#InService}) / measured_throughputApproach 2: Multiply by Arrival Delay
In a stable system without losses, throughput equals the arrival rate ($\lambda = 1 / \text{ARRIVAL_DELAY}$). Therefore, dividing by $\lambda$ is algebraically identical to multiplying by ARRIVAL_DELAY:
(E{#Queue} + E{#InService}) * ARRIVAL_DELAYWhy Approach 2 is Far More Robust
In stationary analysis, both formulations converge to virtually identical numbers:
| Formulation | Stationary Result |
|---|---|
| Approach 1 (divide by measured throughput) | 10.04 min |
| Approach 2 (multiply by arrival delay) | 10.03 min |
However, in transient analysis, Approach 1 breaks down completely:
| Time Instant ($t$) | Approach 1 (divide by measured throughput) | Approach 2 (multiply by arrival delay) |
|---|---|---|
| $t = 25$ | 0.00 | 7.96 min |
| $t = 50$ | 0.00 | 8.80 min |
| $t = 100$ | 0.00 | 10.92 min |
| $t = 200$ | 0.00 | 11.92 min |
Cause of Transient Failure
Transient instantaneous throughput measured from short-lived departure places drops near zero in most replications, triggering zero-division guards. Approach 2 depends only on a known constant parameter (ARRIVAL_DELAY), producing a smooth, mathematically sound curve across all regimes.
Systems with Dropping or Capacity Limits
When a system has finite capacity and drops arriving items when full, only a fraction of arrivals enter the system. The effective arrival rate is scaled by $(1 - P{\text{full}})$:
(E{#Queue} + E{#InService}) * ARRIVAL_DELAY / (1 - P{#Slots = 0})This expression correctly evaluates the average response time for items that were admitted.
Always interpret alongside blocking
Constraining queue capacity shortens the response time of admitted users by shedding load. Always report response time alongside the blocking probability.
Next steps
Learn how to evaluate resource utilization and bottlenecks in Utilization.