Skip to content

Intermittent Unknown error 258 with no obvious cause #1530

Description

@deadwards90

Describe the bug

On occasions we will see the following error

Microsoft.Data.SqlClient.SqlException (0x80131904): Execution Timeout Expired.  The timeout period elapsed prior to completion of the operation or the server is not responding.
 ---> System.ComponentModel.Win32Exception (258): Unknown error 258
   at Microsoft.Data.SqlClient.SqlConnection.OnError(SqlException exception, Boolean breakConnection, Action`1 wrapCloseInAction)
   at Microsoft.Data.SqlClient.TdsParser.ThrowExceptionAndWarning(TdsParserStateObject stateObj, Boolean callerHasConnectionLock, Boolean asyncClose)
   at Microsoft.Data.SqlClient.SqlCommand.InternalEndExecuteReader(IAsyncResult asyncResult, Boolean isInternal, String endMethod)
   at Microsoft.Data.SqlClient.SqlCommand.EndExecuteReaderInternal(IAsyncResult asyncResult)
   at Microsoft.Data.SqlClient.SqlCommand.EndExecuteReaderAsync(IAsyncResult asyncResult)
   at System.Threading.Tasks.TaskFactory`1.FromAsyncCoreLogic(IAsyncResult iar, Func`2 endFunction, Action`1 endAction, Task`1 promise, Boolean requiresSynchronization)

However, SQL Server shows no long running queries and is not using a lot of it's resources during the periods where this happens.

It looks to be more of an intermittent connection issue but we're unable to find any sort of root cause.

To reproduce

We're not sure of the reproduction steps. I've been unable to reproduce this myself by simulating load. From what we can tell this is more likely to happen when the pod is busy (not through just HTTP, but handling events from an external source) but equally it can happen randomly when nothing is really happening on the pod which has caused us quite a substantial amount of confusion.

Expected behavior

Either more information on what the cause might be, or some solution to the issue. I realise the driver might not actually know the issue and it may really be a timeout to it's point of view. We're not entirely sure where the problem lies yet, which is the biggest issue.

Further technical details

Microsoft.Data.SqlClient version: 3.0.1
.NET target: Core 3.1
SQL Server version: Microsoft SQL Azure (RTM) - 12.0.2000.8
Operating system: Docker Container - mcr.microsoft.com/dotnet/aspnet:3.1

Additional context

  • Running in AKS, against Elastic Pools.
  • SQL Server shows no long running queries
  • We sometimes get a TimeoutEvent from the metrics that are collected from the pool. On occasions when we do get them, the error_state will be different.
    • For example, we had one this morning that was 145. We don't know what this means can find no information on what these relate to. I've raised a ticket with the Azure Docs team to look at this. I'll add more onto this when they happen as we've not been keeping track of the error_state codes as we're not sure if they're even relevant.
  • This might be related to this ticket - Execution Timeout Expired Error (258, ReadSniSyncOverAsync) #647
    • However we don't see the ReadSniSyncOverAsync
  • We do have event counter metrics being exported to Prometheus but have found no obvious indicators that something is wrong

Activity

  1. JRahnama commented on Mar 3, 2022

    @JRahnama
    Contributor

    As we have seen before in issue #647 the underlying reason may come from a different problems which may not be entirely drivers fault. any interruption in connectivity a missed port or socket failure could lead to this issue. We cannot say more without having more on the context of application. The most helpful step could be a minimal repro which it is usually impossible to create. However we can capture EventSource traces and see what has gone wrong. I would suggest closing the issue and follow #647.

  2. deadwards90 commented on Mar 4, 2022

    @deadwards90
    Author

    @JRahnama happy to close the issue and add my information to that ticket, but I want to be absolutely sure that the ticket referenced (which I've also included in the original report) is not just for the ReadSniSyncOverAsync errors which are not in our stacktrace.

    In the mean time, we'll hook up the EventSource you linked to gather some more information. Did not realise that was available!


    EDIT: Any suggestions on which event source traces to enable?

  3. JRahnama commented on Mar 4, 2022

    @JRahnama
    Contributor

    @dantheman999301 if you use perfview it will capture all events. make sure you filter them by Microsoft.Data.SqlClient.EventSource name.

  4. deadwards90 commented on Mar 4, 2022

    @deadwards90
    Author

    @JRahnama unfortunately due to the nature of this issue (only showing up in our production environment, which is in Linux Docker in AKS), I think perfview is not going to work for us, as nice as it would be.

    We've managed to wrangle the EventListener so that it will only log on errors. These errors usually show up at least once in the morning for us in a certain service during weekdays so fingers crossed on Monday I should have something for you.

  5. JRahnama commented on Mar 4, 2022

    @JRahnama
    Contributor

    I can happily point you to PerfCollect for Unix machines, but not sure how it works on docker. it captures all events and you can transfer the files to a windows machine and investigate them.

  6. deadwards90 commented on Mar 5, 2022

    @deadwards90
    Author

    @JRahnama good news is, we managed to get the logs (we think).

    Bad news is, there are 200,000 of them. Unfortunately it didn't filter like I thought it would and as it was happening during a busy period and we have all Keywords on, there was a lot to collect.

    We need to work out how we're going to export them from Kibana, and obviously it's not as good as a full dump. Perfcollect might work but we'll need to work out how we're going to hook it up and run it over what is usually a period of an hour without having an impact on our production systems. We might also be able to use dotnet-monitor but it would require some investigation too.

    Let me know if it's of any use.

  7. vincentDAO commented on Mar 10, 2022

    @vincentDAO

    We got same issue 5 days ago even our code was working and didn't change few days ago

  8. JRahnama commented on Mar 10, 2022

    @JRahnama
    Contributor

    @dantheman999301 sorry for the late response we got busy with preview release. Any kind of log that shows or help us to understand where it happens would be helpful.

  9. JRahnama commented on Mar 10, 2022

    @JRahnama
    Contributor

    We got same issue 5 days ago even our code was working and didn't change few days ago

    Same questions applies to your case as well. Any repro or tracing logs would be helpful. I would also suggest investigation network traces as well. that could clarify some of the underlying issues.

  10. oyvost commented on Apr 5, 2022

    @oyvost

    258 timeout is a common exception when the DTU limit is reached on Azure SQL. If on Azure, you can try to monitor the Max DTU percentage and see if it hits the limit.

  11. deadwards90 commented on Apr 20, 2022

    @deadwards90
    Author

    @oyvost In the issue we are facing we can see that the connection is not even made to SQL Server so it's not a throttling issue or similar.

    For example, we had an error about 20 minutes ago and it was utilising 4% of the available DTUs.

    @JRahnama just to get back to you, the logs we thought we had turned out not to be any good, there was a lot of duplication due to the way we wrote the Event Listener. The closest we've got to any answers on this is that we think it's timing out trying to get a connection from the connection pool even though from what we could tell from the event counters there seemed to be plenty of available connections in the pool. We're not overly confident this is the cause but it's something to go off.

    We did see this error once in one of the services that is prone to showing the other error I posted.

    Microsoft.Data.SqlClient.SqlException (0x80131904): A connection was successfully established with the server, but then an error occurred during the pre-login handshake. (provider: TCP Provider, error: 0 - Unknown error 16974573)
     ---> System.ComponentModel.Win32Exception (16974573): Unknown error 16974573
       at Microsoft.Data.ProviderBase.DbConnectionPool.CheckPoolBlockingPeriod(Exception e)
       at Microsoft.Data.ProviderBase.DbConnectionPool.CreateObject(DbConnection owningObject, DbConnectionOptions userOptions, DbConnectionInternal oldConnection)
       at Microsoft.Data.ProviderBase.DbConnectionPool.UserCreateRequest(DbConnection owningObject, DbConnectionOptions userOptions, DbConnectionInternal oldConnection)
       at Microsoft.Data.ProviderBase.DbConnectionPool.TryGetConnection(DbConnection owningObject, UInt32 waitForMultipleObjectsTimeout, Boolean allowCreate, Boolean onlyOneCheckConnection, DbConnectionOptions userOptions, DbConnectionInternal& connection)
       at Microsoft.Data.ProviderBase.DbConnectionPool.WaitForPendingOpen()
    
  12. DLS201 commented on Jul 1, 2022

    @DLS201

    Hello,
    We currently face a similar issue since approximately 1 month, with SQL queries facing Win32Exceptions with code 258 and no obvious cause. The DB itself signals no issue with very low DTU usage.

  13. JRahnama commented on Jul 4, 2022

    @JRahnama
    Contributor

    @DLS201 can you post the stack trace of the exception? have you checked the tcp and socket events/logs?

  14. DLS201 commented on Jul 4, 2022

    @DLS201

    Hi,
    Here is our error message:

    Microsoft.Data.SqlClient.SqlException (0x80131904): Execution Timeout Expired.  The timeout period elapsed prior to completion of the operation or the server is not responding.
    ---> System.ComponentModel.Win32Exception (258): No error information
       at Microsoft.Data.SqlClient.SqlCommand.<>c.<ExecuteDbDataReaderAsync>b__207_0(Task`1 result)
       at System.Threading.Tasks.ContinuationResultTaskFromResultTask`2.InnerInvoke()
       at System.Threading.ExecutionContext.RunInternal(ExecutionContext executionContext, ContextCallback callback, Object state)
    --- End of stack trace from previous location ---
       at System.Threading.Tasks.Task.ExecuteWithThreadLocal(Task& currentTaskSlot, Thread threadPoolThread)
    --- End of stack trace from previous location ---
       at Microsoft.EntityFrameworkCore.Storage.RelationalCommand.ExecuteReaderAsync(RelationalCommandParameterObject parameterObject, CancellationToken cancellationToken)
       at Microsoft.EntityFrameworkCore.Storage.RelationalCommand.ExecuteReaderAsync(RelationalCommandParameterObject parameterObject, CancellationToken cancellationToken)
       at Microsoft.EntityFrameworkCore.Query.Internal.SingleQueryingEnumerable`1.AsyncEnumerator.InitializeReaderAsync(AsyncEnumerator enumerator, CancellationToken cancellationToken)
       at Microsoft.EntityFrameworkCore.Storage.ExecutionStrategy.<>c__DisplayClass33_0`2.<<ExecuteAsync>b__0>d.MoveNext()
    --- End of stack trace from previous location ---
       at Microsoft.EntityFrameworkCore.Storage.ExecutionStrategy.ExecuteImplementationAsync[TState,TResult](Func`4 operation, Func`4 verifySucceeded, TState state, CancellationToken cancellationToken)
       at Microsoft.EntityFrameworkCore.Storage.ExecutionStrategy.ExecuteImplementationAsync[TState,TResult](Func`4 operation, Func`4 verifySucceeded, TState state, CancellationToken cancellationToken)
       at Microsoft.EntityFrameworkCore.Storage.ExecutionStrategy.ExecuteAsync[TState,TResult](TState state, Func`4 operation, Func`4 verifySucceeded, CancellationToken cancellationToken)
       at Microsoft.EntityFrameworkCore.Query.Internal.SingleQueryingEnumerable`1.AsyncEnumerator.MoveNextAsync()
    

    Database in an Azure SQL instance, no error on this side.

  15. mekk1t commented on Jul 19, 2022

    @mekk1t

    @DLS201 Hello! Any progress on this? I'm encountering the same issue, but on three different occassions. I've documented them as a question on Stackoverflow.

  16. 110 remaining items

  17. Will-Bill commented on Aug 18, 2026

    @Will-Bill

    We were hitting this on ASP.NET Core 8 / NHibernate 5.6 / Microsoft.Data.SqlClient 7.0.2, Linux app talking to Azure SQL. Same shape as the original report: intermittent Win32Exception (258) / Execution Timeout Expired, SQL Server quiet, and the query often never arrived at the database. The site would drop in bursts, then recover.

    What actually moved the needle was cutting SqlConnection.Open() churn, not making queries faster.

    NHibernate's default connection.release_mode is after_transaction. In a session-per-request app whose page renders are almost all auto-commit reads, that means check a connection out of the pool for every statement. A typical article page was ~15–20 queries, so ~15–20 Open() / Close() cycles per request. Traces matched that: db.connection count ≈ db.query count.

    We switched HTTP request sessions to ConnectionReleaseMode.OnClose (hold one connection for the request, return it when the session closes):

    sessionFactory.WithOptions()
        .ConnectionReleaseMode(ConnectionReleaseMode.OnClose)
        .OpenSession();

    or

    <property name="connection.release_mode">on_close</property>

    After that, the same pages were ~17 queries on one connection. We left the default in place for background jobs and SignalR, where a session can sit idle and you do not want a pooled connection pinned.

    We also set Connection Lifetime / Load Balance Timeout to 180s (under Azure SNAT's ~4 minute silent idle drop) and ConnectRetryCount=2 on the same deploy, so this is not a perfectly isolated A/B. We did not raise ThreadPool.SetMinThreads.

    Result: no 258 outage periods since (a couple of weeks at time of writing). Article p95 also fell from ~2.7s to ~150ms, which is consistent with Open() being the expensive/failing path, not SQL execution.

    Why this points at pooling / connection management

    For sequential page rendering, peak connections per request is still 1 either way. The change is how many times you go through pool checkout / SqlConnection.Open() per request — and therefore how often you risk a stale pooled socket, sp_reset_connection, a physical reconnect/TLS handshake, or the Linux SNI wait that surfaces as error 258.

    If reducing connection opens (not query count, not Max Pool Size, not min threads) stops the 258s, the failure is in pool checkout / Open / SNI wait, not command execution on SQL Server.

    EF Core default is the same idea. A scoped DbContext does not hold a SQL connection for the request. EF opens just before each operation and closes afterwards (back to the ADO.NET pool). AddDbContextPool only reuses the DbContext object; it does not change that. So 15 queries on one context still typically means 15 pool checkouts, unless you have an explicit transaction or you opened the connection yourself. Equivalent experiment:

    await context.Database.OpenConnectionAsync();
    try
    {
        // all queries in this request reuse one pooled connection
    }
    finally
    {
        await context.Database.CloseConnectionAsync();
    }

    Caveat: holding the connection for the request keeps it across any non-SQL work in the same request (outbound HTTP, etc.). Do not do this on long-lived idle sessions.

  18. federico-paganini commented on Aug 25, 2026

    @federico-paganini

    This issue is a catch-all, and I think that is why four years of platform-vs-platform comparison never closed it. We now have two different deterministic repros that both end in Unknown error 258, plus a way to tell them apart from the timings you already have. Everything below is measured or read from source; both repros run anywhere Docker does.

    @MichelZ asked on 2025-11-12 for a Docker-based reproduction. Cause B below is exactly that, and it also answers the question about the v7 async path.


    Cause A: the pool hands out a dead socket

    A middlebox (SNAT/egress on ACA, App Service, NAT Gateway, corporate firewalls) silently drops an idle outbound flow — no RST, no FIN. The socket stays ESTABLISHED on the client, and the pool's checkout-time liveness check cannot tell the difference.

    This thread has had enough theory, so both halves of that are measured rather than read. The pool does check — WaitHandleDbConnectionPool.GetConnection calls IsConnectionAlive(), and has since at least release/6.1. The check is SniTcpHandle.CheckConnection, which reduces to !_socket.Connected || (_socket.Poll(100, SelectMode.SelectRead) && _socket.Available == 0). Evaluated character for character on three real sockets against a SQL Server container:

    socket state Connected Poll(100, SelectRead) Available check reports
    healthy, idle True False 0 SNI_SUCCESS — alive
    flow black-holed True False 0 SNI_SUCCESS — alive
    peer closed it (control) True True 0 SNI_ERROR — dead, evicted

    The first two rows are identical in every observable. The predicate is not wrong; there is no local signal that separates a black-holed socket from a healthy idle one. The third row is the control — the check does exactly what it was designed for.

    The same contrast through SqlClient itself, with ClientConnectionId as the identity of the physical connection: warm one pooled connection, return it to the pool, disturb it, then take one from that pool again.

    control: KILL <spid> (peer closes) subject: flow black-holed
    Open() 10 ms 1 ms
    ClientConnectionId different — the pool evicted the corpse same — the pool kept it
    outcome SELECT 1 -> 1 in 0.0 s Number=-2 + Win32 258 after 20.0 s

    The control is what makes this conclusive: the pool evicts a corpse it can see, so "no eviction" in the subject row is blindness, not the absence of a check.

    To be precise about a distinction this thread has blurred: @MichelZ's SNAT port exhaustion (2022-09) was correctly ruled out. This is @Malcolm-Stewart's 2024-02-14 point one layer down — the idle drop: no ports run out; an idle flow is discarded silently.

    Why it surfaces as a timeout that says 258. Once the command is in flight, TCP keepalive is out of play — it only runs on idle sockets — so the kernel's retransmission budget governs. On Linux (tcp_retries2=15) that budget is ~15 minutes, so your CommandTimeout always wins. SqlClient then sends a TDS Attention and waits AttentionTimeoutSeconds = 5 (TdsParserStateObject.cs) for an ACK that can never arrive. That predicts failure at CommandTimeout + 5 s exactly:

    CommandTimeout 17 s 30 s 43 s 61 s 120 s
    Fails at 22.0 s 35.0 s 48.0 s 66.0 s 125.0 s

    Independent of ConnectRetryCount (tested 0/1/2). EF Core deliberately does not retry -2, which is why it reaches the user. This is also why "the query never arrived at the database" (@deadwards90) and why the server shows nothing — it never left the client. It is the shape @jdudleyie reported in 2022-12-17: ~2-minute operations in Application Insights against a 30-second client timeout, with no long-running query server-side.

    Why moving to Windows "fixes" it — it doesn't, it changes which timer wins. We could not test a Windows host, so we changed the one variable instead, on the same repro:

    tcp_retries2 Wins the race Result
    15 (Linux default) your CommandTimeout Number=-2 + Win32 258 at CT+5 s — not retried by EF, user-visible
    5 (Windows-like budget) the kernel Number=0, "A transport-level error has occurred", at 13.1 s — retryable, the driver reconnects and nobody notices
    15, with CommandTimeout=1200 s the kernel same transport-level error at 938.7 s

    That last row is the control: give Linux a timeout longer than its retransmission budget and it produces exactly the same retryable error Windows produces. The driver behaves identically on both platforms — only the two clocks differ, and both are configurable. This also suggests #2103 ("transport-level error") and this issue may be one bug seen from two kernel configurations.

    The fix, demonstrated. Same machine, same simulated middlebox, CommandTimeout=30:

    Open() Outcome
    6.1.4 1 ms — serves the dead connection fails at 35.0 s, Number=-2 + 258
    7.1.0-preview2 + Connection Idle Timeout=10 + UseLegacyIdleTimeoutBehavior=false 13 ms — discards it, opens a new one query OK in 0.0 s

    Connection Idle Timeout (#4295, 7.1) is the first setting that evicts a connection for sitting idle. Two caveats: it is gated behind that AppContext switch, which defaults to true — with the default the keyword parses and does nothing. And it is time-based rather than liveness-based, which is exactly why it works here — no local check can detect this. Per #4295 it evicts both on retrieval and on a periodic pass (idleTimeout/2 cadence for WaitHandleDbConnectionPool). The remaining race is between the value you configure and the middlebox's own timer, so set it comfortably below your egress's idle window.

    Connection Lifetime / Load Balance Timeout is not a substitute. Measured on 6.1.4: it is evaluated on return, never at checkout, so a connection left idle in the pool is handed straight back out.


    Cause B: closing a reader early, drained under the command timeout

    This is @DerPeit's cause (2024-05-21), which we reproduced faithfully. The network is healthy here and the query does reach the server — everything about it is different from cause A except the error text.

    ExecuteReader returns as soon as the first packet lands. Closing the reader before the result set is consumed drains the remaining rows, and that drain runs under the command timeout that is still armed. It cannot finish, the timeout fires, and on Linux the read timeout surfaces as Number=-2 wrapping Win32Exception: Unknown error 258.

    It has never been fixed. Same repro, CommandTimeout=1, seven versions across three majors:

    5.1.5 5.2.3 6.0.2 6.1.4 6.1.6 7.0.2 7.1.0-preview2
    258 @ 1.2 s 258 @ 1.2 s 258 @ 1.2 s 258 @ 1.2 s 258 @ 1.2 s 258 @ 1.2 s 258 @ 1.2 s

    That row covers the net8.0 assembly. MDS ships a separate net9.0 one, so we re-ran 6.0.2 through 7.1.0-preview2 against it as well, on .NET 9 and .NET 10 (a .NET 10 app resolves to the net9.0 asset, verified from the loaded assembly's TargetFrameworkAttribute, and on a different base image). Every cell identical.

    @MichelZ, to answer your question directly: the new 7.x async path does not change this. With UseCompatibilityAsyncBehaviour=false and UseCompatibilityProcessSni=false, on both 7.0.2 and 7.1.0-preview2, and on both assemblies, the result is identical. Encrypt=True, MultipleActiveResultSets=True, and whether any rows were read before closing make no difference either.

    Why some callers see it and others never do

    The axis is the closing verb, and it is orthogonal to sync vs async:

    Close() Dispose()
    sync throws 258 swallowed
    async throws 258 swallowed

    SqlDataReader.Dispose(bool) catches and discards every SqlException from Close(), leaving only an EventSource trace. So a plain using (var reader = cmd.ExecuteReader()) hides this, while an explicit reader.Close() throws.

    EF Core users always see it because RelationalDataReader calls the throwing verb on purpose in both teardown paths — _reader.Close() in Dispose, await _reader.CloseAsync() in DisposeAsync, each annotated // can throw in EF's own source. SqlDataReader does not override CloseAsync, so the base implementation runs Close().

    Confirmed end to end with @DerPeit's own foreach + break:

    stack sync foreach await foreach
    EF 8.0.5 + MDS 5.1.5 (his era) 258 @ 1.5 s 258 @ 1.5 s
    EF 9.0.0 + MDS 7.1.0-preview2 (today) 258 @ 1.5 s 258 @ 1.5 s

    Dapper is not affected, which answers the question @DerPeit left open. With buffered: false and breaking out early it completes cleanly in 0.2 s, because Dapper already calls cmd?.Cancel() during teardown (three sites in SqlMapper.cs) — the same thing his interceptor does, built in.

    The workaround in this thread is incomplete

    This matters for everyone who applied it. @DerPeit's interceptor overrides DataReaderClosing, which is the sync hook only:

    interceptor sync foreach await foreach
    DataReaderClosing only (as published) ok @ 0.5 s 258 @ 1.5 s
    DataReaderClosing + DataReaderClosingAsync ok @ 0.5 s ok @ 0.5 s

    Stable over 3 runs per cell, identical on the 2024 stack and today's, and on .NET 8, 9 and 10. Async call sites — most of a modern ASP.NET application — were never covered. @fr4gles, this is very likely what you hit on 2025-11-12 ("surely help a lot ... but it did not solve entirely"). The complete version:

    public class CommandCancelingInterceptor : DbCommandInterceptor
    {
        public override InterceptionResult DataReaderClosing(
            DbCommand command, DataReaderClosingEventData eventData, InterceptionResult result)
        {
            command.Cancel();
            return base.DataReaderClosing(command, eventData, result);
        }
    
        public override async ValueTask<InterceptionResult> DataReaderClosingAsync(
            DbCommand command, DataReaderClosingEventData eventData, InterceptionResult result)
        {
            command.Cancel();
            return await base.DataReaderClosingAsync(command, eventData, result);
        }
    }

    Two things we could not confirm

    @DerPeit reported two further variants. The one where a per-row transform throws does reproduce, and the 258 from teardown replaces the original exception entirely — the caller never learns why their code actually failed. That one is arguably the worst of the three, because it destroys the evidence.

    The cancellation-only variant did not reproduce for us. Running his code as written — the per-row Transform that awaits Task.Yield() and returns normally — we always got a clean TaskCanceledException, on both EF 8.0.5 + MDS 5.1.5 and EF 9.0.0 + MDS 7.1.0-preview2, sweeping the cancellation moment from 100 ms to 5 s (nine points), at CommandTimeout of both 1 s and 30 s, and with the thread pool floored at one worker. Fourteen runs, no 258, the exception always tracking the cancellation instant.

    The mechanism explains why: EF passes the token down to ReadAsync, SqlClient sends an Attention, the server abandons the query, and the subsequent teardown has nothing left to drain. That is the binding @DerPeit suspected was too late — it does not appear to be. We are not claiming this was fixed, since it is equally clean on his era's stack; our environment differs from his in at least one respect we cannot rule out (he ran Server=localhost with Integrated Security=True on Windows, which may not even be a TCP connection).

    We also checked whether an abandoned drain poisons the pool, since that would explain the "random 258 on an idle server" reports. It does not. After both the throwing and the swallowing path the connection returns to the pool, the next borrow gets the same ClientConnectionId, and SELECT 1 answers in 0.00 s. Cause B only ever bills the caller that closed the reader.


    Telling the two apart from timings you already have

    Set CommandTimeout to an odd value, say 47 s, and look at where the 258s land.

    • CommandTimeout + a small, latency-proportional tail → cause B. The tail is a couple of network round-trips: +0.2 s on a LAN, +0.5 s at 50 ms RTT, +1.5 s at 200 ms RTT (measured through a latency proxy).
    • CommandTimeout + 5.0 s, constant → cause A. That 5 s is fixed regardless of CommandTimeout and of RTT, because it is AttentionTimeoutSeconds waiting for an ACK from a socket with nothing alive behind it.

    With tracing, cause A has a second signature: compare each request's wall-clock against the sum of its dependency spans. It produces ~1 ms of actual SQL and tens of seconds unaccounted for, because tracing instruments command execution, not connection acquisition. Slow connection establishment looks the same in that view, so it separates mechanism classes, not every case.


    Both repros are small and self-contained. Cause A drops a single TCP flow with iptables (needs --cap-add=NET_ADMIN, no host root, no Azure); cause B needs only a SQL Server container. Happy to post either here, or to contribute them as fixtures to the pool reliability test framework offered in #3667, so this class of report becomes reproducible in CI instead of only in production. Glad to run any variant on request — the whole matrix above takes minutes.

    For the pool redesign (#3356), the gap behind cause A: there is no liveness validation at checkout, so a connection killed by anything external is served exactly once as if it were healthy. Idle-based eviction narrows that window; test-on-borrow (JDBC isValid() / HikariCP semantics) is what closes it.

  19. federico-paganini commented on Aug 25, 2026

    @federico-paganini

    A follow-up on cause B from my earlier comment. I first framed this as a documentation gap; having since read #288, that framing was wrong, so here it is narrowed to what I think actually stands.

    What was already established. In #288 (2019) @Wraith2 described this mechanism precisely: "output parameters are filled in as they are encountered in the result stream ... if the execution is cancelled/closed before the return values are passed back they won't be filled in." @David-Engel confirmed there that not cancelling on close is deliberate and common to all the SQL Server drivers, with @roji's reasoning that cancelling a batched command could skip later statements. None of that is in question, and the doc wording that came out of #288 is accurate.

    What I measured. A stored procedure that returns a large result set and then sets an output parameter, with no rows read before the reader is torn down:

    scenario @out (procedure sets 42) RETURN (procedure returns 7) exception
    drain completes — TOP 10, CommandTimeout=30 42 7 none
    drain aborted — 20M rows, CommandTimeout=1 — Dispose() null null none
    drain aborted — 20M rows, CommandTimeout=1 — Close() null null Number=-2

    Same on 5.1.5, 6.1.4 and 7.1.0-preview2. The first row is the control showing the procedure and the harness are sound.

    The narrow point. Rows two and three differ in one thing only: whether the caller is told. SqlDataReader.Dispose(bool) catches the SqlException from Close() and leaves an EventSource trace behind. So using (var reader = cmd.ExecuteReader()) returns normally after an aborted drain, and the caller then reads a null output parameter with nothing available to explain why.

    I do not think anyone decided that. It reads like two individually sound decisions composing: "do not cancel on close" (settled in #288, for good reasons) and "do not throw from Dispose" (the general .NET guideline). Between them sits a path that yields wrong values silently, and I could not find an open issue covering it.

    For anyone hitting this today, the practical answer is the one @roji gives in dotnet/efcore#24857 — async plus a cancellation token. I can confirm it behaves correctly here: cancelling mid-stream produces a clean TaskCanceledException rather than a 258, on both EF 8.0.5 + MDS 5.1.5 and EF 9.0.0 + MDS 7.1.0-preview2, across cancellation points from 100 ms to 5 s.

    Whether the swallow is worth revisiting, and whether this belongs in its own issue rather than as a note in this thread, seems like a call for someone on the team rather than mine to make.

  20. fr4gles commented on Aug 25, 2026

    @fr4gles

    I can confirm that mentioned workaround works for most of our cases but we also introduced retry logic all over our database access layer to mitigate this error.

    It's kinda ok now but it still rarely hits us randomly out of nowhere - retry logic handles it.
    Very frustrating TBH.

  21. federico-paganini commented on Aug 26, 2026

    @federico-paganini

    One more measured observation, this time about the error text rather than the mechanism.

    SqlException builds its inner exception as new Win32Exception(errorCollection[0].Win32ErrorCode), and the codes stored there are Windows codes — TdsEnums.SNI_WAIT_TIMEOUT = 258 among them. On Unix, Win32Exception resolves the number through strerror, so the number is read as an errno. Measured on .NET 8, 9 and 10 (Debian 12 and Ubuntu 24.04 — identical in all three):

    code message on Linux on Windows
    258 Unknown error 258 WAIT_TIMEOUT — "The wait operation timed out."
    53 Invalid request descriptor ERROR_BAD_NETPATH — "The network path was not found."
    10054 Unknown error 10054 WSAECONNRESET
    110 Connection timed out 110 is also a valid errno (ETIMEDOUT)

    Two shapes come out of that. Where the number has no errno, the message is "Unknown error N" — 258 is this issue's title. Where it collides with an unrelated errno, the message is that errno's text instead; 53 is the code from #1773.

    I noticed #1773 was closed as resolved by #3461, which aligned the Numbers across platforms and added the SQL_ConnectTimeout string. The inner-exception text above still reproduces on 6.1.6 and on 7.1.0-preview2, the newest published build. I do not know whether that part was considered in scope there, so I am reporting the measurement rather than assuming anything was missed.

    If someone on the team is able to review this and confirm it, that would be interesting.

  22. deadwards90 commented on Sep 3, 2026

    @deadwards90
    Author

    @federico-paganini Nice repro! Wish I'd had an AI to throw at this back in the day, although looks like our guess of something funky with the connection pool wasn't too far off.

    I just want to check part of your explanation though:

    Independent of ConnectRetryCount (tested 0/1/2). EF Core deliberately does not retry -2, which is why it reaches the user. This is also why "the query never arrived at the database" (@deadwards90) and why the server shows nothing — it never left the client. It is the shape @jdudleyie reported in 2022-12-17: ~2-minute operations in Application Insights against a 30-second client timeout, with no long-running query server-side.

    We weren't using EF Core, we were using Dapper. I assume the explanation still holds in this case though.

  23. federico-paganini commented on Sep 3, 2026

    @federico-paganini

    @deadwards90 Thanks! And yes — it has been an excellent tool for this. Getting here took some fairly wild detours: per-flow netfilter black-holes, the kernel's TCP knobs (tcp_retries2) to change which timer wins, reading /proc/net/tcp to find the one socket to kill.

    On Dapper: it holds, and rather than assume it we ran it — same per-flow black-hole, same 30 s CommandTimeout, Dapper (QueryFirstAsync) vs raw SqlCommand, on 3.0.1 (your version at the time) and 6.1.4:

    MDS mapper Open() fails at Number inner
    3.0.1 SqlCommand 1 ms, reused the dead connection 35.0 s −2 Win32 258
    3.0.1 Dapper 1 ms, reused the dead connection 35.0 s −2 Win32 258
    6.1.4 SqlCommand 1 ms, reused the dead connection 35.0 s −2 Win32 258
    6.1.4 Dapper 1 ms, reused the dead connection 35.0 s −2 Win32 258

    Indistinguishable. The mechanism sits below any mapper: the pool hands out the socket, the first write is black-holed, the command timer fires, and the attention the driver then sends is black-holed too — AttentionTimeoutSeconds = 5 in TdsParserStateObject, hence CommandTimeout + 5 s exactly.

    The EF Core sentence was only about why nothing above SqlClient swallows the -2 in an EF app: SqlServerTransientExceptionDetector has //case -2: commented out, with the reason stated right above it — "This exception can be thrown even if the operation completed successfully, so it's safer to let the application fail." Dapper has no SqlException retry layer anywhere in the package (its only "retry" is an ArgumentException flags fallback for SQLite), so with Dapper the -2 reaches the caller with one layer fewer. Your original trace — Execution Timeout Expired, Win32Exception (258) in EndExecuteReaderAsync — is this shape, and it fits what you saw in 2022: the query never reached the server, so the server had nothing to report and DTU was irrelevant.

    On the pool: the locus was right, the mechanism differs from the 2022 guess. Checkout does not time out — it returns the dead connection in 1 ms, because the liveness check it runs is a local Socket.Poll, which cannot see a silently dropped flow. The command then times out on it.

  24. mdaigle commented on Sep 3, 2026

    @mdaigle
    Contributor

    Thanks @federico-paganini, this is a great analysis!

    For Cause A:
    I'm glad that Connection Idle Timeout can be helpful here. Like you mentioned, there's no way to tell if a flow has been silently dropped other than to run a full round trip to the server when checking a connection out from the pool. That would be prohibitively expensive on higher latency connections. I've thought about other options like a background liveness/keep-alive process, but it's also heavy-weight and would drive a lot of additional traffic through the server.

    Connection Idle Timeout will be available in the 7.1.0 release (it's also in 7.1.0-preview3). I decided to gate it behind an app context switch because it has the potential to cause unexpected behavior. It was developed as part of the new pool design which is also opt-in. If community feedback is favorable, I'll enable it by default along with the new pool in the 8.0 major version release.

    For Cause B:
    An improved error message would be helpful here. Things are "working as expected", but the message is completely opaque and doesn't provide any guidance on how to avoid the issue. I'll open a separate issue for this.

    It looks like we also need a separate issue for this:

    @DerPeit reported two further variants. The one where a per-row transform throws does reproduce, and the 258 from teardown replaces the original exception entirely — the caller never learns why their code actually failed. That one is arguably the worst of the three, because it destroys the evidence.

    I feel it would be better to surface all the exceptions via an AggregateException.

  25. federico-paganini commented on Sep 3, 2026

    @federico-paganini

    @mdaigle Thank you. The 8.0 timing is consistent with the one cost we measured of enforcing idle expiry: with Min Pool Size, retainees are expired at checkout only and the floor never re-warms on its own, so the first N callers after a long idle gap each pay a cold open. It is in the #4581 body.

    Two pointers for the new issues, both already measured:

    Glad to have been of help. We'll follow #4638 and #4639.

  26. edwardneal commented on Sep 4, 2026

    @edwardneal
    Contributor

    I've been thinking about cause A, and I think this can partially be fixed by making KeepAliveTime, KeepAliveCount and possible KeepAliveInterval into connection string parameters. Both the JDBC driver and SqlClient already use these to detect broken idle connections.

    Doing this allows users to circumvent the OS-specific tcp_retries2 parameter by changing the TCP keepalive timers (and telling the OS to clean up the connection's local end if TCP discovers that it's died.) Operating at the TCP layer means that the SQL Server won't need to execute a query, and enabling these parameters to be changed by users allows them to compensate for the idle connection timeouts of various middleboxes.

    It'd need a few pieces of guidance in the documentation:

    • KeepAliveTime + (KeepAliveInterval * KeepAliveCount) should be less than or equal to the amount of time the application is willing to wait for a stale connection to be detected. One factor involved here may be the middlebox's TCP session timeout (for example, the default for Azure's NAT gateway is four minutes).
    • If the formula's result is too low, connections in the connection pool would be unnecessarily doomed.
    • If the formula's result is too high, we find ourselves in the same situation: a middlebox can silently drop an idle TCP connection.
    • Depending upon the middlebox, the simple act of enabling TCP keepalives could change the result - some middleboxes might preserve the TCP connection because they're seeing regular traffic across it.

    These new options could map to the TCP_KEEPIDLE, TCP_KEEPCNT and TCP_KEEPINTVL TCP parameters on the socket (plus the SO_KEEPALIVE parameter) in Linux, and the TcpKeepAliveTime, TcpKeepAliveRetryCount, TcpKeepAliveInterval and KeepAlive options on a .NET socket.

    Importantly, these three connection string parameters would be part of the connection pool key, and would be applied at the point of opening the physical connection.

    If this helps to address cause A, Azure SQL Gateway has a known quirk which needs to be considered. I think this can be overcome by making sure that the existing connection pool's default Connection Idle Timeout is less than 30 minutes when connected to an Azure SQL endpoint.


    While this should hopefully help to address the first part of cause A (the connection pool will hopefully be made less likely to issue dead sockets) we'd need to add an extra piece of work to adjust the unacknowledged transmission timeout. On Linux, we could perhaps look to adjust the TCP_USER_TIMEOUT option. TCP_MAXRT initially appears to achieve a similar idea within Windows for our purposes.

    Given that the objective is to make sure that we detect a silently killed TCP connection within a period of time, I'm inclined to say that we should set either of these parameters to KeepAliveTime + (KeepAliveInterval * KeepAliveCount). This would maintain a similar failure-detection period irrespective of whether the TCP connection dies silently within or outwith the connection pool.


    Edit: KeepAliveTime and KeepAliveCount are not valid connection string parameters in mssql-jdbc, I misread the documentation. Instead, mssql-jdbc sets KeepAliveTime / TCP_KEEPIDLE to 30 seconds and KeepAliveInterval to 1 second (link). So does SqlClient (link) I also added a second point: while these should prevent the pool from handing out dead sockets, TCP_USER_TIMEOUT/TCP_MAXRT may be helpful when a connection has been issued by the pool and is being used by the client. Finally: the documentation I looked at for KeepAliveCount for Windows was wrong. Newer versions of Windows do support this after all.

  27. priyankatiwari08 commented on Sep 17, 2026

    @priyankatiwari08
    Contributor

    /triage

  28. github-actions commented on Sep 17, 2026

    @github-actions

    test placeholder - to be removed

    Generated by SqlClient Issue Auto-Triage for #1530 · copilot · auto · 110.1 AIC · ⌖ 13.2 AIC · ⊞ 11.9K · ◷

  29. x-itm commented on Sep 18, 2026

    @x-itm

    The timing can help separate the two cases. If Open() succeeds quickly, SQL shows no corresponding request, and failure occurs at roughly CommandTimeout (often plus the teardown/attention interval), that is consistent with a pooled connection whose outbound flow was silently dropped—not a slow server query.

    For affected workloads, correlate connection checkout/idle duration with failures and compare against any NAT/firewall idle timeout. The new pool's Connection Idle Timeout option (noted above for 7.1) is worth testing in a controlled rollout: it trades some cold opens after idle periods for avoiding reuse of long-idle sockets. Keep retries targeted to transient failures and ensure commands are idempotent before retrying.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Area\Managed SNIIssues that are targeted to the Managed SNI codebase.Performance 📈Issues that are targeted to performance improvements.

    Type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions