Signal Waits: Telling CPU Pressure From Slow Disks

High processor use and slow reads can both make a query feel stuck. Signal waits measure the time spent waiting for CPU after a resource becomes ready, helping you separate those delays.

Three swimmers waiting in line on a jetty for a swimmer in a red cap to leave the only diving board

Split Signal Waits From Resource Time

A task first waits for its resource, then waits to run again after the resource is available. The latter interval is signal wait time. The total wait counter includes both portions. Subtract the signal portion to estimate the resource portion for completed recorded waits.

I read those portions together rather than labeling every large wait total as a slow resource. A disk-related wait can include time waiting to return to the processor. If the resource finished promptly but the ready queue remained crowded, the storage device deserves a more careful judgment.

The counters in sys.dm_os_wait_stats accumulate since startup or a reset. A lifetime proportion mixes busy periods, quiet periods, maintenance, and older incidents. Collect two snapshots around the affected workload. Keep the server startup time and capture interval with them. Yesterday's average is a poor witness to today's queue.

Establish a Bounded Counter Snapshot

The following script stores a first snapshot in temporary tables within your session. Run the later block after a representative observation interval. Do not clear shared wait statistics to simplify the exercise. Another monitoring process depends on those counters too.

SELECT sqlserver_start_time AS StartupTime
INTO #SignalStart FROM sys.dm_os_sys_info;
SELECT wait_type, waiting_tasks_count, wait_time_ms,
       signal_wait_time_ms
INTO #SignalBefore FROM sys.dm_os_wait_stats;
SELECT SYSDATETIME() AS CaptureStarted;

Record the capture time and workload separately. Use the same session so the temporary tables remain available. The query does not wait for a fixed duration or manufacture load. Choose an interval that contains the actual operation you are investigating, then capture the second state.

Compute Signal Waits From Valid Deltas

Reject the comparison if the instance restarted. Also reject counter decreases, which indicate a reset or otherwise invalid pairing. Keep filters explicit. Removing idle and background waits changes the denominator, so a percentage is meaningful only alongside its selected population.

IF EXISTS (SELECT 1 FROM #SignalStart AS b
           CROSS JOIN sys.dm_os_sys_info AS e
           WHERE b.StartupTime <> e.sqlserver_start_time)
    THROW 51000, 'The instance restarted during capture.', 1;
SELECT e.wait_type,
       e.wait_time_ms - b.wait_time_ms AS TotalWaitMs,
       e.signal_wait_time_ms - b.signal_wait_time_ms AS SignalWaitMs
INTO #SignalDelta
FROM sys.dm_os_wait_stats AS e
JOIN #SignalBefore AS b ON b.wait_type = e.wait_type;
IF EXISTS (SELECT 1 FROM #SignalDelta
           WHERE TotalWaitMs < 0 OR SignalWaitMs < 0)
    THROW 51001, 'Wait counters decreased during capture.', 1;
SELECT SUM(TotalWaitMs) AS TotalWaitMs,
       SUM(SignalWaitMs) AS SignalWaitMs,
       SUM(TotalWaitMs - SignalWaitMs) AS ResourceWaitMs,
       100.0 * SUM(SignalWaitMs) /
           NULLIF(SUM(TotalWaitMs), 0) AS SignalSharePercent
FROM #SignalDelta
WHERE wait_type NOT IN
    ('SLEEP_TASK', 'LAZYWRITER_SLEEP', 'SQLTRACE_BUFFER_FLUSH',
     'BROKER_TASK_STOP', 'XE_TIMER_EVENT', 'XE_DISPATCHER_WAIT');

Review that exclusion list for your monitoring policy. It illustrates a bounded calculation rather than a universal benign-wait catalog. A reset followed by enough new activity to exceed the old values can evade simple decrease detection. Coordinate captures with known reset operations when that risk matters.

Keep Absolute Work Beside the Percentage

A high proportion over a tiny wait total has different significance from sustained scheduling delay over a busy interval. Return actual totals and execution activity alongside the share. Compare application latency, throughput, and processor use. There is no single percentage that proves every server needs more CPU.

Wait counters describe accumulated task waiting, not elapsed wall-clock time for the whole server. Parallel tasks can contribute concurrently. Do not divide total waits by observation duration and call the result a utilization percentage. The numerator and denominator describe different populations and can produce misleading conclusions.

I inspect the largest contributing wait types after the aggregate. Their resource and signal portions explain whether one workload dominates the sample. Include the capture filters in the report so the next person can reproduce the arithmetic rather than debating a percentage without its inputs.

SELECT TOP (20) wait_type, TotalWaitMs, SignalWaitMs,
       TotalWaitMs - SignalWaitMs AS ResourceWaitMs
FROM #SignalDelta
WHERE TotalWaitMs > 0
ORDER BY TotalWaitMs DESC;
Two parts of every completed wait: a diagram about the signal waits

Confirm Signal Waits With Scheduler Queues

Scheduler ready queues provide current evidence of work awaiting CPU. Read visible online schedulers and repeat the sample during the slow operation. Correlate the pattern with Windows processor counters. A single sample after the incident has ended cannot reconstruct the earlier queue.

SELECT scheduler_id, cpu_id, runnable_tasks_count,
       current_tasks_count, active_workers_count
FROM sys.dm_os_schedulers
WHERE status = 'VISIBLE ONLINE'
ORDER BY runnable_tasks_count DESC;

Broad sustained queues and high processor use support a capacity or excess-work investigation. Uneven queues need attention to scheduler distribution and serial workload limits. Check the virtual host too when applicable. Available guest CPU does not always mean uninterrupted access to physical execution time.

Compare Storage Latency Over the Same Period

File I/O counters also accumulate, so take a first snapshot before the same observation interval. The next block calculates per-file read and write latency from changes. File names identify where the observed activity occurred. Read and write totals need their own denominators.

SELECT database_id, file_id, num_of_reads, num_of_writes,
       io_stall_read_ms, io_stall_write_ms
INTO #FileBefore
FROM sys.dm_io_virtual_file_stats(NULL, NULL);

Run this before the interval, alongside the first wait snapshot. Afterward, execute the next query. Use a stable interval without a restart or counter reset, and inspect negative deltas. Newly created files without an earlier row are excluded by the join and need separate reporting.

SELECT DB_NAME(e.database_id) AS DatabaseName, f.name AS FileName,
       (e.io_stall_read_ms - b.io_stall_read_ms) * 1.0 /
           NULLIF(e.num_of_reads - b.num_of_reads, 0) AS AverageReadMs,
       (e.io_stall_write_ms - b.io_stall_write_ms) * 1.0 /
           NULLIF(e.num_of_writes - b.num_of_writes, 0) AS AverageWriteMs
FROM sys.dm_io_virtual_file_stats(NULL, NULL) AS e
JOIN #FileBefore AS b
  ON b.database_id = e.database_id AND b.file_id = e.file_id
JOIN sys.master_files AS f
  ON f.database_id = e.database_id AND f.file_id = e.file_id
WHERE e.num_of_reads >= b.num_of_reads
  AND e.num_of_writes >= b.num_of_writes
  AND e.io_stall_read_ms >= b.io_stall_read_ms
  AND e.io_stall_write_ms >= b.io_stall_write_ms;

Investigate Necessary Work Before Adding Capacity

A processor queue can grow because queries repeatedly perform avoidable work. Inspect the expensive statements for large scans, repeated scalar calculations, and joins that multiply rows. Parameter-sensitive plans and incorrect estimates can change that work between executions. Improving the access path can reduce scheduling pressure without changing processor capacity.

Storage pressure also needs a query explanation. A healthy device serving an unnecessary flood of reads still deserves an indexing or query review. Conversely, a reasonable access path does not erase genuine device latency. Keep logical reads, physical I/O activity, and measured file stalls visible together. They answer related but different questions about the same workload.

Avoid diagnosing CPU from parallel synchronization waits alone. Coordination can reflect uneven worker progress or another slow resource. Review the actual plan and worker behavior. Restricting parallelism changes resource sharing, but it does not automatically remove the original cause. Test that change against the whole workload before retaining it.

Combine the Clues Before Choosing a Fix

Which evidence agrees with the users' slow period? Rising scheduling delay with sustained ready queues supports CPU investigation. Large resource portions and elevated file-latency deltas support storage or excessive I/O investigation. Both pressures can exist together, so avoid forcing every incident into one category.

Signal waits help locate the delayed execution stage. Review query CPU, reads, plans, and concurrency to decide whether capacity or unnecessary work caused it. Signal waits become useful when paired with bounded captures and current queue evidence. Measure again after one justified change, using comparable workload conditions.

Related reading on this blog: Detecting CPU Pressure with Wait Statistics and CPU Scheduler Waiting On Disk.

What the signal share proves: a checklist on the signal waits

A wait percentage is not a bottleneck diagnosis, it is one part of an interval-based resource investigation.

Published by Pinal Dave on SQLAuthority. More of my work at pinaldave.com.

SQL CPU, SQL DMV, SQL Performance, SQL Wait Stats
Previous Post
When In-Memory OLTP Helps and When It Does Not
Next Post
Unstable Queries: Finding Wide Duration Swings in Query Store

Related Posts

2 Comments. Leave new

Leave a Reply

Your email address will not be published. Required fields are marked *

Fill out this field
Fill out this field
Please enter a valid email address.