A DBA on call needs a short path from an alarm to useful evidence. Establish the impact, confirm the instance, and narrow the failure before making a change.

Establish What Has Actually Failed
Ask which action is failing, which users are affected, and when the symptoms began. Record the time zone with the timestamp. An unavailable application does not automatically mean the database engine has stopped.
Try the approved connection path and record the exact error. If you cannot connect, check the service, network path, and recent platform events through the operations process. Do not spend the first ten minutes editing connection settings at random.
SELECT SYSDATETIMEOFFSET() AS observed_at,
@@SERVERNAME AS server_name,
DB_NAME() AS database_name,
SERVERPROPERTY('ProductVersion') AS product_version;
SELECT name, state_desc, user_access_desc
FROM sys.databases;Once connected, verify that this is the affected instance. Read database states before assuming every database is available. Keep the incident owner informed with facts and the next check.
Look for Work Waiting Behind Other Work
Inspect active requests for blocking and wait types. A large blocking chain can make the entire application look unavailable. The session at its head might be sleeping while holding an open transaction.
SELECT r.session_id, r.status, r.command,
r.wait_type, r.wait_time, r.blocking_session_id,
r.cpu_time, r.total_elapsed_time,
s.login_name, s.host_name, s.program_name
FROM sys.dm_exec_requests AS r
JOIN sys.dm_exec_sessions AS s ON s.session_id = r.session_id
WHERE r.session_id <> @@SPID
ORDER BY r.total_elapsed_time DESC;
SELECT session_id, status, open_transaction_count,
login_name, host_name, program_name
FROM sys.dm_exec_sessions
WHERE is_user_process = 1 AND open_transaction_count > 0;Identify the owner and purpose of a blocking transaction before ending it. Cancellation can lead to rollback, which also takes time and resources. A restart is an expensive substitute for understanding the queue.
Check Capacity in the Right Place
A full data volume, log volume, or backup destination creates different problems. Inspect the volume hosting each database file and check operating-system evidence too. The following query covers volumes visible through database files, not every destination.
SELECT DISTINCT vs.volume_mount_point,
vs.total_bytes / 1073741824.0 AS total_gb,
vs.available_bytes / 1073741824.0 AS available_gb
FROM sys.master_files AS mf
CROSS APPLY sys.dm_os_volume_stats(mf.database_id, mf.file_id) AS vs;
SELECT name, recovery_model_desc, log_reuse_wait_desc
FROM sys.databases
WHERE database_id > 4;Log reuse waits can explain why space inside a log is not becoming reusable. Investigate the reported reason before proposing a shrink. Do not delete database files or backup chains to make an alarm disappear.
If approved retention cleanup is needed, identify exactly what remains recoverable afterward. Expanding capacity may be the safer immediate action. Follow the established owner and change process rather than improvising file removal.
Read Recent Job Outcomes and Errors
A failed overnight load can explain missing data without any engine outage. Check recent Agent job outcomes and open the detailed step history for the relevant failure. The final job message alone may omit the useful error.
SELECT TOP (20) j.name, h.run_date, h.run_time,
h.run_status, h.message
FROM msdb.dbo.sysjobhistory AS h
JOIN msdb.dbo.sysjobs AS j ON j.job_id = h.job_id
WHERE h.step_id = 0 AND h.run_status <> 1
ORDER BY h.instance_id DESC;This history can include canceled or other non-successful outcomes, not only failures. Jobs still running need a separate activity check. Express does not include SQL Server Agent, so check the external scheduler used there.
EXEC master.dbo.sp_readerrorlog 0, 1, N'Error';Searching for Error is a starting filter, not a complete incident record. Read neighboring messages and earlier log files when necessary. Correlate timestamps with application and Windows events.
Ask What Changed
Check recent deployments, maintenance, credential changes, failovers, and workload shifts. A change near the incident is a lead rather than proof of causation. Compare it with the evidence already collected.
Use an existing monitoring baseline or Query Store when available. Current snapshots cannot reconstruct every earlier condition. Be explicit when the required history was never captured.
Choose One Action and Verify Its Effect
State the proposed action, expected result, and reversal path before carrying it out. Use the approved incident authority for disruptive operations. Then verify the original user action rather than merely watching a dashboard turn green.
Keep a timeline with commands, observations, and decisions. Hand over unresolved risks and any temporary workaround. Three in the morning is a poor time to trust memory, even when the coffee is confident.
An on-call checklist is not a repair script, it is a way to make the next decision safely.
This post was rewritten from scratch in September 2026. The original, published on 2011-10-26, was a short announcement about something that no longer exists. The address is the same, the subject is now something worth keeping.
Published by Pinal Dave on SQLAuthority. More of my work at pinaldave.com.
Discover more from SQL Authority with Pinal Dave
Subscribe to get the latest posts sent to your email.





10 Comments. Leave new
Pinal is a real inspiration to everyone in the SQL Server Community and a true supporter of new software solutions such as our innovative data integration software.
Having Pinal and Rick at our booth #300 for an incredible book signing event was a highlight at this year’s PASS Summit.
Thanks,
Michael
Once again, Congratulations Pinal! You are making India proud.
You are role model for DBAs like me.
Congratulations sir!
Proud of you sir ! May this be the first of many more to come.
Raguram
Congratulations on the book signing, neat achievement.
Congratulations Pinal.
Congratulations Pinal! This is truly a commendable achievement, which I am sure will inspire others. :)
congrats pinal
I have brought Joes 2 Pros % Book set from IndiaPlaza on 26 Otober 2012. How to get free SQl wais stats..
Read the details about how to receive books here – http://blog.sqlauthority.com/2012/10/15/sql-server-free-print-book-on-sql-server-joes-2-pros-kit/