Repository navigation
Intermittent Unknown error 258 with no obvious cause #1530
Description
Activity
As we have seen before in issue #647 the underlying reason may come from a different problems which may not be entirely drivers fault. any interruption in connectivity a missed port or socket failure could lead to this issue. We cannot say more without having more on the context of application. The most helpful step could be a minimal repro which it is usually impossible to create. However we can capture EventSource traces and see what has gone wrong. I would suggest closing the issue and follow #647.
@JRahnama happy to close the issue and add my information to that ticket, but I want to be absolutely sure that the ticket referenced (which I've also included in the original report) is not just for the
ReadSniSyncOverAsyncerrors which are not in our stacktrace.In the mean time, we'll hook up the EventSource you linked to gather some more information. Did not realise that was available!
EDIT: Any suggestions on which event source traces to enable?
@dantheman999301 if you use perfview it will capture all events. make sure you filter them by
Microsoft.Data.SqlClient.EventSourcename.@JRahnama unfortunately due to the nature of this issue (only showing up in our production environment, which is in Linux Docker in AKS), I think perfview is not going to work for us, as nice as it would be.
We've managed to wrangle the
EventListenerso that it will only log on errors. These errors usually show up at least once in the morning for us in a certain service during weekdays so fingers crossed on Monday I should have something for you.I can happily point you to PerfCollect for Unix machines, but not sure how it works on docker. it captures all events and you can transfer the files to a windows machine and investigate them.
@JRahnama good news is, we managed to get the logs (we think).
Bad news is, there are 200,000 of them. Unfortunately it didn't filter like I thought it would and as it was happening during a busy period and we have all Keywords on, there was a lot to collect.
We need to work out how we're going to export them from Kibana, and obviously it's not as good as a full dump. Perfcollect might work but we'll need to work out how we're going to hook it up and run it over what is usually a period of an hour without having an impact on our production systems. We might also be able to use
dotnet-monitorbut it would require some investigation too.Let me know if it's of any use.
We got same issue 5 days ago even our code was working and didn't change few days ago
@dantheman999301 sorry for the late response we got busy with preview release. Any kind of log that shows or help us to understand where it happens would be helpful.
We got same issue 5 days ago even our code was working and didn't change few days ago
Same questions applies to your case as well. Any repro or tracing logs would be helpful. I would also suggest investigation network traces as well. that could clarify some of the underlying issues.
258 timeout is a common exception when the DTU limit is reached on Azure SQL. If on Azure, you can try to monitor the Max DTU percentage and see if it hits the limit.
@oyvost In the issue we are facing we can see that the connection is not even made to SQL Server so it's not a throttling issue or similar.
For example, we had an error about 20 minutes ago and it was utilising 4% of the available DTUs.
@JRahnama just to get back to you, the logs we thought we had turned out not to be any good, there was a lot of duplication due to the way we wrote the Event Listener. The closest we've got to any answers on this is that we think it's timing out trying to get a connection from the connection pool even though from what we could tell from the event counters there seemed to be plenty of available connections in the pool. We're not overly confident this is the cause but it's something to go off.
We did see this error once in one of the services that is prone to showing the other error I posted.
Microsoft.Data.SqlClient.SqlException (0x80131904): A connection was successfully established with the server, but then an error occurred during the pre-login handshake. (provider: TCP Provider, error: 0 - Unknown error 16974573) ---> System.ComponentModel.Win32Exception (16974573): Unknown error 16974573 at Microsoft.Data.ProviderBase.DbConnectionPool.CheckPoolBlockingPeriod(Exception e) at Microsoft.Data.ProviderBase.DbConnectionPool.CreateObject(DbConnection owningObject, DbConnectionOptions userOptions, DbConnectionInternal oldConnection) at Microsoft.Data.ProviderBase.DbConnectionPool.UserCreateRequest(DbConnection owningObject, DbConnectionOptions userOptions, DbConnectionInternal oldConnection) at Microsoft.Data.ProviderBase.DbConnectionPool.TryGetConnection(DbConnection owningObject, UInt32 waitForMultipleObjectsTimeout, Boolean allowCreate, Boolean onlyOneCheckConnection, DbConnectionOptions userOptions, DbConnectionInternal& connection) at Microsoft.Data.ProviderBase.DbConnectionPool.WaitForPendingOpen()Hello,
We currently face a similar issue since approximately 1 month, with SQL queries facing Win32Exceptions with code 258 and no obvious cause. The DB itself signals no issue with very low DTU usage.Reacted by Maly Lemire@DLS201 can you post the stack trace of the exception? have you checked the tcp and socket events/logs?
Hi,
Here is our error message:Microsoft.Data.SqlClient.SqlException (0x80131904): Execution Timeout Expired. The timeout period elapsed prior to completion of the operation or the server is not responding. ---> System.ComponentModel.Win32Exception (258): No error information at Microsoft.Data.SqlClient.SqlCommand.<>c.<ExecuteDbDataReaderAsync>b__207_0(Task`1 result) at System.Threading.Tasks.ContinuationResultTaskFromResultTask`2.InnerInvoke() at System.Threading.ExecutionContext.RunInternal(ExecutionContext executionContext, ContextCallback callback, Object state) --- End of stack trace from previous location --- at System.Threading.Tasks.Task.ExecuteWithThreadLocal(Task& currentTaskSlot, Thread threadPoolThread) --- End of stack trace from previous location --- at Microsoft.EntityFrameworkCore.Storage.RelationalCommand.ExecuteReaderAsync(RelationalCommandParameterObject parameterObject, CancellationToken cancellationToken) at Microsoft.EntityFrameworkCore.Storage.RelationalCommand.ExecuteReaderAsync(RelationalCommandParameterObject parameterObject, CancellationToken cancellationToken) at Microsoft.EntityFrameworkCore.Query.Internal.SingleQueryingEnumerable`1.AsyncEnumerator.InitializeReaderAsync(AsyncEnumerator enumerator, CancellationToken cancellationToken) at Microsoft.EntityFrameworkCore.Storage.ExecutionStrategy.<>c__DisplayClass33_0`2.<<ExecuteAsync>b__0>d.MoveNext() --- End of stack trace from previous location --- at Microsoft.EntityFrameworkCore.Storage.ExecutionStrategy.ExecuteImplementationAsync[TState,TResult](Func`4 operation, Func`4 verifySucceeded, TState state, CancellationToken cancellationToken) at Microsoft.EntityFrameworkCore.Storage.ExecutionStrategy.ExecuteImplementationAsync[TState,TResult](Func`4 operation, Func`4 verifySucceeded, TState state, CancellationToken cancellationToken) at Microsoft.EntityFrameworkCore.Storage.ExecutionStrategy.ExecuteAsync[TState,TResult](TState state, Func`4 operation, Func`4 verifySucceeded, CancellationToken cancellationToken) at Microsoft.EntityFrameworkCore.Query.Internal.SingleQueryingEnumerable`1.AsyncEnumerator.MoveNextAsync()Database in an Azure SQL instance, no error on this side.
@DLS201 Hello! Any progress on this? I'm encountering the same issue, but on three different occassions. I've documented them as a question on Stackoverflow.
110 remaining items
We were hitting this on ASP.NET Core 8 / NHibernate 5.6 / Microsoft.Data.SqlClient 7.0.2, Linux app talking to Azure SQL. Same shape as the original report: intermittent
Win32Exception (258)/Execution Timeout Expired, SQL Server quiet, and the query often never arrived at the database. The site would drop in bursts, then recover.What actually moved the needle was cutting
SqlConnection.Open()churn, not making queries faster.NHibernate's default
connection.release_modeisafter_transaction. In a session-per-request app whose page renders are almost all auto-commit reads, that means check a connection out of the pool for every statement. A typical article page was ~15–20 queries, so ~15–20Open()/Close()cycles per request. Traces matched that:db.connectioncount ≈db.querycount.We switched HTTP request sessions to
ConnectionReleaseMode.OnClose(hold one connection for the request, return it when the session closes):sessionFactory.WithOptions() .ConnectionReleaseMode(ConnectionReleaseMode.OnClose) .OpenSession();
or
<property name="connection.release_mode">on_close</property>
After that, the same pages were ~17 queries on one connection. We left the default in place for background jobs and SignalR, where a session can sit idle and you do not want a pooled connection pinned.
We also set
Connection Lifetime/Load Balance Timeoutto 180s (under Azure SNAT's ~4 minute silent idle drop) andConnectRetryCount=2on the same deploy, so this is not a perfectly isolated A/B. We did not raiseThreadPool.SetMinThreads.Result: no 258 outage periods since (a couple of weeks at time of writing). Article p95 also fell from ~2.7s to ~150ms, which is consistent with
Open()being the expensive/failing path, not SQL execution.Why this points at pooling / connection management
For sequential page rendering, peak connections per request is still 1 either way. The change is how many times you go through pool checkout /
SqlConnection.Open()per request — and therefore how often you risk a stale pooled socket,sp_reset_connection, a physical reconnect/TLS handshake, or the Linux SNI wait that surfaces as error 258.If reducing connection opens (not query count, not Max Pool Size, not min threads) stops the 258s, the failure is in pool checkout / Open / SNI wait, not command execution on SQL Server.
EF Core default is the same idea. A scoped
DbContextdoes not hold a SQL connection for the request. EF opens just before each operation and closes afterwards (back to the ADO.NET pool).AddDbContextPoolonly reuses theDbContextobject; it does not change that. So 15 queries on one context still typically means 15 pool checkouts, unless you have an explicit transaction or you opened the connection yourself. Equivalent experiment:await context.Database.OpenConnectionAsync(); try { // all queries in this request reuse one pooled connection } finally { await context.Database.CloseConnectionAsync(); }
Caveat: holding the connection for the request keeps it across any non-SQL work in the same request (outbound HTTP, etc.). Do not do this on long-lived idle sessions.
Reacted by FranekThis issue is a catch-all, and I think that is why four years of platform-vs-platform comparison never closed it. We now have two different deterministic repros that both end in
Unknown error 258, plus a way to tell them apart from the timings you already have. Everything below is measured or read from source; both repros run anywhere Docker does.@MichelZ asked on 2025-11-12 for a Docker-based reproduction. Cause B below is exactly that, and it also answers the question about the v7 async path.
Cause A: the pool hands out a dead socket
A middlebox (SNAT/egress on ACA, App Service, NAT Gateway, corporate firewalls) silently drops an idle outbound flow — no RST, no FIN. The socket stays
ESTABLISHEDon the client, and the pool's checkout-time liveness check cannot tell the difference.This thread has had enough theory, so both halves of that are measured rather than read. The pool does check —
WaitHandleDbConnectionPool.GetConnectioncallsIsConnectionAlive(), and has since at leastrelease/6.1. The check isSniTcpHandle.CheckConnection, which reduces to!_socket.Connected || (_socket.Poll(100, SelectMode.SelectRead) && _socket.Available == 0). Evaluated character for character on three real sockets against a SQL Server container:socket state ConnectedPoll(100, SelectRead)Availablecheck reports healthy, idle True False 0 SNI_SUCCESS— aliveflow black-holed True False 0 SNI_SUCCESS— alivepeer closed it (control) True True 0 SNI_ERROR— dead, evictedThe first two rows are identical in every observable. The predicate is not wrong; there is no local signal that separates a black-holed socket from a healthy idle one. The third row is the control — the check does exactly what it was designed for.
The same contrast through SqlClient itself, with
ClientConnectionIdas the identity of the physical connection: warm one pooled connection, return it to the pool, disturb it, then take one from that pool again.control: KILL <spid>(peer closes)subject: flow black-holed Open()10 ms 1 ms ClientConnectionIddifferent — the pool evicted the corpse same — the pool kept it outcome SELECT 1-> 1 in 0.0 sNumber=-2+ Win32 258 after 20.0 sThe control is what makes this conclusive: the pool evicts a corpse it can see, so "no eviction" in the subject row is blindness, not the absence of a check.
To be precise about a distinction this thread has blurred: @MichelZ's SNAT port exhaustion (2022-09) was correctly ruled out. This is @Malcolm-Stewart's 2024-02-14 point one layer down — the idle drop: no ports run out; an idle flow is discarded silently.
Why it surfaces as a timeout that says 258. Once the command is in flight, TCP keepalive is out of play — it only runs on idle sockets — so the kernel's retransmission budget governs. On Linux (
tcp_retries2=15) that budget is ~15 minutes, so yourCommandTimeoutalways wins. SqlClient then sends a TDS Attention and waitsAttentionTimeoutSeconds = 5(TdsParserStateObject.cs) for an ACK that can never arrive. That predicts failure atCommandTimeout+ 5 s exactly:CommandTimeout 17 s 30 s 43 s 61 s 120 s Fails at 22.0 s 35.0 s 48.0 s 66.0 s 125.0 s Independent of
ConnectRetryCount(tested 0/1/2). EF Core deliberately does not retry-2, which is why it reaches the user. This is also why "the query never arrived at the database" (@deadwards90) and why the server shows nothing — it never left the client. It is the shape @jdudleyie reported in 2022-12-17: ~2-minute operations in Application Insights against a 30-second client timeout, with no long-running query server-side.Why moving to Windows "fixes" it — it doesn't, it changes which timer wins. We could not test a Windows host, so we changed the one variable instead, on the same repro:
tcp_retries2Wins the race Result 15 (Linux default) your CommandTimeoutNumber=-2+ Win32 258 at CT+5 s — not retried by EF, user-visible5 (Windows-like budget) the kernel Number=0, "A transport-level error has occurred", at 13.1 s — retryable, the driver reconnects and nobody notices15, with CommandTimeout=1200 sthe kernel same transport-level error at 938.7 s That last row is the control: give Linux a timeout longer than its retransmission budget and it produces exactly the same retryable error Windows produces. The driver behaves identically on both platforms — only the two clocks differ, and both are configurable. This also suggests #2103 ("transport-level error") and this issue may be one bug seen from two kernel configurations.
The fix, demonstrated. Same machine, same simulated middlebox,
CommandTimeout=30:Open()Outcome 6.1.41 ms — serves the dead connection fails at 35.0 s, Number=-2+ 2587.1.0-preview2+Connection Idle Timeout=10+UseLegacyIdleTimeoutBehavior=false13 ms — discards it, opens a new one query OK in 0.0 s Connection Idle Timeout(#4295, 7.1) is the first setting that evicts a connection for sitting idle. Two caveats: it is gated behind thatAppContextswitch, which defaults totrue— with the default the keyword parses and does nothing. And it is time-based rather than liveness-based, which is exactly why it works here — no local check can detect this. Per #4295 it evicts both on retrieval and on a periodic pass (idleTimeout/2cadence forWaitHandleDbConnectionPool). The remaining race is between the value you configure and the middlebox's own timer, so set it comfortably below your egress's idle window.Connection Lifetime/Load Balance Timeoutis not a substitute. Measured on 6.1.4: it is evaluated on return, never at checkout, so a connection left idle in the pool is handed straight back out.
Cause B: closing a reader early, drained under the command timeout
This is @DerPeit's cause (2024-05-21), which we reproduced faithfully. The network is healthy here and the query does reach the server — everything about it is different from cause A except the error text.
ExecuteReaderreturns as soon as the first packet lands. Closing the reader before the result set is consumed drains the remaining rows, and that drain runs under the command timeout that is still armed. It cannot finish, the timeout fires, and on Linux the read timeout surfaces asNumber=-2wrappingWin32Exception: Unknown error 258.It has never been fixed. Same repro,
CommandTimeout=1, seven versions across three majors:5.1.5 5.2.3 6.0.2 6.1.4 6.1.6 7.0.2 7.1.0-preview2 258 @ 1.2 s 258 @ 1.2 s 258 @ 1.2 s 258 @ 1.2 s 258 @ 1.2 s 258 @ 1.2 s 258 @ 1.2 s That row covers the
net8.0assembly. MDS ships a separatenet9.0one, so we re-ran 6.0.2 through 7.1.0-preview2 against it as well, on .NET 9 and .NET 10 (a .NET 10 app resolves to thenet9.0asset, verified from the loaded assembly'sTargetFrameworkAttribute, and on a different base image). Every cell identical.@MichelZ, to answer your question directly: the new 7.x async path does not change this. With
UseCompatibilityAsyncBehaviour=falseandUseCompatibilityProcessSni=false, on both 7.0.2 and 7.1.0-preview2, and on both assemblies, the result is identical.Encrypt=True,MultipleActiveResultSets=True, and whether any rows were read before closing make no difference either.Why some callers see it and others never do
The axis is the closing verb, and it is orthogonal to sync vs async:
Close()Dispose()sync throws 258 swallowed async throws 258 swallowed SqlDataReader.Dispose(bool)catches and discards everySqlExceptionfromClose(), leaving only an EventSource trace. So a plainusing (var reader = cmd.ExecuteReader())hides this, while an explicitreader.Close()throws.EF Core users always see it because
RelationalDataReadercalls the throwing verb on purpose in both teardown paths —_reader.Close()inDispose,await _reader.CloseAsync()inDisposeAsync, each annotated// can throwin EF's own source.SqlDataReaderdoes not overrideCloseAsync, so the base implementation runsClose().Confirmed end to end with @DerPeit's own
foreach+break:stack sync foreachawait foreachEF 8.0.5 + MDS 5.1.5 (his era) 258 @ 1.5 s 258 @ 1.5 s EF 9.0.0 + MDS 7.1.0-preview2 (today) 258 @ 1.5 s 258 @ 1.5 s Dapper is not affected, which answers the question @DerPeit left open. With
buffered: falseand breaking out early it completes cleanly in 0.2 s, because Dapper already callscmd?.Cancel()during teardown (three sites inSqlMapper.cs) — the same thing his interceptor does, built in.The workaround in this thread is incomplete
This matters for everyone who applied it. @DerPeit's interceptor overrides
DataReaderClosing, which is the sync hook only:interceptor sync foreachawait foreachDataReaderClosingonly (as published)ok @ 0.5 s 258 @ 1.5 s DataReaderClosing+DataReaderClosingAsyncok @ 0.5 s ok @ 0.5 s Stable over 3 runs per cell, identical on the 2024 stack and today's, and on .NET 8, 9 and 10. Async call sites — most of a modern ASP.NET application — were never covered. @fr4gles, this is very likely what you hit on 2025-11-12 ("surely help a lot ... but it did not solve entirely"). The complete version:
public class CommandCancelingInterceptor : DbCommandInterceptor { public override InterceptionResult DataReaderClosing( DbCommand command, DataReaderClosingEventData eventData, InterceptionResult result) { command.Cancel(); return base.DataReaderClosing(command, eventData, result); } public override async ValueTask<InterceptionResult> DataReaderClosingAsync( DbCommand command, DataReaderClosingEventData eventData, InterceptionResult result) { command.Cancel(); return await base.DataReaderClosingAsync(command, eventData, result); } }
Two things we could not confirm
@DerPeit reported two further variants. The one where a per-row transform throws does reproduce, and the 258 from teardown replaces the original exception entirely — the caller never learns why their code actually failed. That one is arguably the worst of the three, because it destroys the evidence.
The cancellation-only variant did not reproduce for us. Running his code as written — the per-row
Transformthat awaitsTask.Yield()and returns normally — we always got a cleanTaskCanceledException, on both EF 8.0.5 + MDS 5.1.5 and EF 9.0.0 + MDS 7.1.0-preview2, sweeping the cancellation moment from 100 ms to 5 s (nine points), atCommandTimeoutof both 1 s and 30 s, and with the thread pool floored at one worker. Fourteen runs, no 258, the exception always tracking the cancellation instant.The mechanism explains why: EF passes the token down to
ReadAsync, SqlClient sends an Attention, the server abandons the query, and the subsequent teardown has nothing left to drain. That is the binding @DerPeit suspected was too late — it does not appear to be. We are not claiming this was fixed, since it is equally clean on his era's stack; our environment differs from his in at least one respect we cannot rule out (he ranServer=localhostwithIntegrated Security=Trueon Windows, which may not even be a TCP connection).We also checked whether an abandoned drain poisons the pool, since that would explain the "random 258 on an idle server" reports. It does not. After both the throwing and the swallowing path the connection returns to the pool, the next borrow gets the same
ClientConnectionId, andSELECT 1answers in 0.00 s. Cause B only ever bills the caller that closed the reader.
Telling the two apart from timings you already have
Set
CommandTimeoutto an odd value, say 47 s, and look at where the 258s land.CommandTimeout+ a small, latency-proportional tail → cause B. The tail is a couple of network round-trips: +0.2 s on a LAN, +0.5 s at 50 ms RTT, +1.5 s at 200 ms RTT (measured through a latency proxy).CommandTimeout+ 5.0 s, constant → cause A. That 5 s is fixed regardless ofCommandTimeoutand of RTT, because it isAttentionTimeoutSecondswaiting for an ACK from a socket with nothing alive behind it.
With tracing, cause A has a second signature: compare each request's wall-clock against the sum of its dependency spans. It produces ~1 ms of actual SQL and tens of seconds unaccounted for, because tracing instruments command execution, not connection acquisition. Slow connection establishment looks the same in that view, so it separates mechanism classes, not every case.
Both repros are small and self-contained. Cause A drops a single TCP flow with
iptables(needs--cap-add=NET_ADMIN, no host root, no Azure); cause B needs only a SQL Server container. Happy to post either here, or to contribute them as fixtures to the pool reliability test framework offered in #3667, so this class of report becomes reproducible in CI instead of only in production. Glad to run any variant on request — the whole matrix above takes minutes.For the pool redesign (#3356), the gap behind cause A: there is no liveness validation at checkout, so a connection killed by anything external is served exactly once as if it were healthy. Idle-based eviction narrows that window; test-on-borrow (JDBC
isValid()/ HikariCP semantics) is what closes it.Reacted by Michael Einsiedler, Michel Zehnder, Franek, Johan Kronberg, Daniel Edwards, Karl, José Arrarte and Marcus NilssonReacted by FranekReacted by Franek and Malcolm DaigleA follow-up on cause B from my earlier comment. I first framed this as a documentation gap; having since read #288, that framing was wrong, so here it is narrowed to what I think actually stands.
What was already established. In #288 (2019) @Wraith2 described this mechanism precisely: "output parameters are filled in as they are encountered in the result stream ... if the execution is cancelled/closed before the return values are passed back they won't be filled in." @David-Engel confirmed there that not cancelling on close is deliberate and common to all the SQL Server drivers, with @roji's reasoning that cancelling a batched command could skip later statements. None of that is in question, and the doc wording that came out of #288 is accurate.
What I measured. A stored procedure that returns a large result set and then sets an output parameter, with no rows read before the reader is torn down:
scenario @out(procedure sets 42)RETURN (procedure returns 7) exception drain completes — TOP 10,CommandTimeout=3042 7 none drain aborted — 20M rows, CommandTimeout=1—Dispose()null null none drain aborted — 20M rows, CommandTimeout=1—Close()null null Number=-2Same on 5.1.5, 6.1.4 and 7.1.0-preview2. The first row is the control showing the procedure and the harness are sound.
The narrow point. Rows two and three differ in one thing only: whether the caller is told.
SqlDataReader.Dispose(bool)catches theSqlExceptionfromClose()and leaves an EventSource trace behind. Sousing (var reader = cmd.ExecuteReader())returns normally after an aborted drain, and the caller then reads a null output parameter with nothing available to explain why.I do not think anyone decided that. It reads like two individually sound decisions composing: "do not cancel on close" (settled in #288, for good reasons) and "do not throw from
Dispose" (the general .NET guideline). Between them sits a path that yields wrong values silently, and I could not find an open issue covering it.For anyone hitting this today, the practical answer is the one @roji gives in dotnet/efcore#24857 — async plus a cancellation token. I can confirm it behaves correctly here: cancelling mid-stream produces a clean
TaskCanceledExceptionrather than a 258, on both EF 8.0.5 + MDS 5.1.5 and EF 9.0.0 + MDS 7.1.0-preview2, across cancellation points from 100 ms to 5 s.Whether the swallow is worth revisiting, and whether this belongs in its own issue rather than as a note in this thread, seems like a call for someone on the team rather than mine to make.
Reacted by Franek, David Engel and Marcus NilssonI can confirm that mentioned workaround works for most of our cases but we also introduced retry logic all over our database access layer to mitigate this error.
It's kinda ok now but it still rarely hits us randomly out of nowhere - retry logic handles it.
Very frustrating TBH.One more measured observation, this time about the error text rather than the mechanism.
SqlExceptionbuilds its inner exception asnew Win32Exception(errorCollection[0].Win32ErrorCode), and the codes stored there are Windows codes —TdsEnums.SNI_WAIT_TIMEOUT = 258among them. On Unix,Win32Exceptionresolves the number throughstrerror, so the number is read as an errno. Measured on .NET 8, 9 and 10 (Debian 12 and Ubuntu 24.04 — identical in all three):code message on Linux on Windows 258 Unknown error 258WAIT_TIMEOUT— "The wait operation timed out."53 Invalid request descriptorERROR_BAD_NETPATH— "The network path was not found."10054 Unknown error 10054WSAECONNRESET110 Connection timed out110 is also a valid errno ( ETIMEDOUT)Two shapes come out of that. Where the number has no errno, the message is "Unknown error N" — 258 is this issue's title. Where it collides with an unrelated errno, the message is that errno's text instead; 53 is the code from #1773.
I noticed #1773 was closed as resolved by #3461, which aligned the
Numbers across platforms and added theSQL_ConnectTimeoutstring. The inner-exception text above still reproduces on 6.1.6 and on 7.1.0-preview2, the newest published build. I do not know whether that part was considered in scope there, so I am reporting the measurement rather than assuming anything was missed.If someone on the team is able to review this and confirm it, that would be interesting.
@federico-paganini Nice repro! Wish I'd had an AI to throw at this back in the day, although looks like our guess of something funky with the connection pool wasn't too far off.
I just want to check part of your explanation though:
Independent of ConnectRetryCount (tested 0/1/2). EF Core deliberately does not retry -2, which is why it reaches the user. This is also why "the query never arrived at the database" (@deadwards90) and why the server shows nothing — it never left the client. It is the shape @jdudleyie reported in 2022-12-17: ~2-minute operations in Application Insights against a 30-second client timeout, with no long-running query server-side.
We weren't using EF Core, we were using Dapper. I assume the explanation still holds in this case though.
@deadwards90 Thanks! And yes — it has been an excellent tool for this. Getting here took some fairly wild detours: per-flow netfilter black-holes, the kernel's TCP knobs (
tcp_retries2) to change which timer wins, reading/proc/net/tcpto find the one socket to kill.On Dapper: it holds, and rather than assume it we ran it — same per-flow black-hole, same 30 s
CommandTimeout, Dapper (QueryFirstAsync) vs rawSqlCommand, on 3.0.1 (your version at the time) and 6.1.4:MDS mapper Open()fails at Numberinner 3.0.1 SqlCommand 1 ms, reused the dead connection 35.0 s −2 Win32 258 3.0.1 Dapper 1 ms, reused the dead connection 35.0 s −2 Win32 258 6.1.4 SqlCommand 1 ms, reused the dead connection 35.0 s −2 Win32 258 6.1.4 Dapper 1 ms, reused the dead connection 35.0 s −2 Win32 258 Indistinguishable. The mechanism sits below any mapper: the pool hands out the socket, the first write is black-holed, the command timer fires, and the attention the driver then sends is black-holed too —
AttentionTimeoutSeconds = 5inTdsParserStateObject, henceCommandTimeout + 5 sexactly.The EF Core sentence was only about why nothing above SqlClient swallows the
-2in an EF app:SqlServerTransientExceptionDetectorhas//case -2:commented out, with the reason stated right above it — "This exception can be thrown even if the operation completed successfully, so it's safer to let the application fail." Dapper has noSqlExceptionretry layer anywhere in the package (its only "retry" is anArgumentExceptionflags fallback for SQLite), so with Dapper the-2reaches the caller with one layer fewer. Your original trace —Execution Timeout Expired,Win32Exception (258)inEndExecuteReaderAsync— is this shape, and it fits what you saw in 2022: the query never reached the server, so the server had nothing to report and DTU was irrelevant.On the pool: the locus was right, the mechanism differs from the 2022 guess. Checkout does not time out — it returns the dead connection in 1 ms, because the liveness check it runs is a local
Socket.Poll, which cannot see a silently dropped flow. The command then times out on it.Reacted by Daniel EdwardsThanks @federico-paganini, this is a great analysis!
For Cause A:
I'm glad that Connection Idle Timeout can be helpful here. Like you mentioned, there's no way to tell if a flow has been silently dropped other than to run a full round trip to the server when checking a connection out from the pool. That would be prohibitively expensive on higher latency connections. I've thought about other options like a background liveness/keep-alive process, but it's also heavy-weight and would drive a lot of additional traffic through the server.Connection Idle Timeout will be available in the 7.1.0 release (it's also in 7.1.0-preview3). I decided to gate it behind an app context switch because it has the potential to cause unexpected behavior. It was developed as part of the new pool design which is also opt-in. If community feedback is favorable, I'll enable it by default along with the new pool in the 8.0 major version release.
For Cause B:
An improved error message would be helpful here. Things are "working as expected", but the message is completely opaque and doesn't provide any guidance on how to avoid the issue. I'll open a separate issue for this.It looks like we also need a separate issue for this:
@DerPeit reported two further variants. The one where a per-row transform throws does reproduce, and the 258 from teardown replaces the original exception entirely — the caller never learns why their code actually failed. That one is arguably the worst of the three, because it destroys the evidence.
I feel it would be better to surface all the exceptions via an AggregateException.
@mdaigle Thank you. The 8.0 timing is consistent with the one cost we measured of enforcing idle expiry: with
Min Pool Size, retainees are expired at checkout only and the floor never re-warms on its own, so the first N callers after a long idle gap each pay a cold open. It is in the #4581 body.Two pointers for the new issues, both already measured:
- For Improve the timeout error when closing an incompletely consumed SqlDataReader #4639, the inner text is a separate artifact from the message — and it was fixed on one path but not this one. Align SqlException Numbers across platforms #3461 (which closed SQLException.Number property differs on Linux OS and Windows. #1773) now throws
new Win32Exception(TdsEnums.SNI_WAIT_TIMEOUT, Strings.SQL_ConnectTimeout)when a connection attempt times out, so that 258 reads "The connection attempt timed out." on every OS. The 258 in this issue comes from the command/teardown timeout, whereSqlExceptionstill buildsnew Win32Exception(Win32ErrorCode)with no message, and on Unix the runtime resolves the Windows code throughstrerror_r: "Unknown error 258" (and 53 becomes the plausible-and-false "Invalid request descriptor"). Reproduced after Align SqlException Numbers across platforms #3461 on 6.1.6 and 7.1.0-preview2, identical on .NET 8, 9 and 10 — details in this comment. It is the same code path for cause A and cause B. - The reproductions behind both issues are at https://github.com/federico-paganini/sqlclient-1530-repros — one mode per variant (
sync-close/async-close/poison-sync/poison-async/ef/ef-async/dapper/outparams-*), with the package version selectable at build time so the era-exact combinations can be re-run, plus the cause-A harness with the Dapper mode and the pool liveness probe.
- For Improve the timeout error when closing an incompletely consumed SqlDataReader #4639, the inner text is a separate artifact from the message — and it was fixed on one path but not this one. Align SqlException Numbers across platforms #3461 (which closed SQLException.Number property differs on Linux OS and Windows. #1773) now throws
I've been thinking about cause A, and I think this can partially be fixed by making
KeepAliveTime,KeepAliveCountand possibleKeepAliveIntervalinto connection string parameters. Both the JDBC driver and SqlClient already use these to detect broken idle connections.Doing this allows users to circumvent the OS-specific
tcp_retries2parameter by changing the TCP keepalive timers (and telling the OS to clean up the connection's local end if TCP discovers that it's died.) Operating at the TCP layer means that the SQL Server won't need to execute a query, and enabling these parameters to be changed by users allows them to compensate for the idle connection timeouts of various middleboxes.It'd need a few pieces of guidance in the documentation:
KeepAliveTime + (KeepAliveInterval * KeepAliveCount)should be less than or equal to the amount of time the application is willing to wait for a stale connection to be detected. One factor involved here may be the middlebox's TCP session timeout (for example, the default for Azure's NAT gateway is four minutes).- If the formula's result is too low, connections in the connection pool would be unnecessarily doomed.
- If the formula's result is too high, we find ourselves in the same situation: a middlebox can silently drop an idle TCP connection.
- Depending upon the middlebox, the simple act of enabling TCP keepalives could change the result - some middleboxes might preserve the TCP connection because they're seeing regular traffic across it.
These new options could map to the
TCP_KEEPIDLE,TCP_KEEPCNTandTCP_KEEPINTVLTCP parameters on the socket (plus theSO_KEEPALIVEparameter) in Linux, and theTcpKeepAliveTime,TcpKeepAliveRetryCount,TcpKeepAliveIntervalandKeepAliveoptions on a .NET socket.Importantly, these three connection string parameters would be part of the connection pool key, and would be applied at the point of opening the physical connection.
If this helps to address cause A, Azure SQL Gateway has a known quirk which needs to be considered. I think this can be overcome by making sure that the existing connection pool's default Connection Idle Timeout is less than 30 minutes when connected to an Azure SQL endpoint.
While this should hopefully help to address the first part of cause A (the connection pool will hopefully be made less likely to issue dead sockets) we'd need to add an extra piece of work to adjust the unacknowledged transmission timeout. On Linux, we could perhaps look to adjust the
TCP_USER_TIMEOUToption.TCP_MAXRTinitially appears to achieve a similar idea within Windows for our purposes.Given that the objective is to make sure that we detect a silently killed TCP connection within a period of time, I'm inclined to say that we should set either of these parameters to
KeepAliveTime + (KeepAliveInterval * KeepAliveCount). This would maintain a similar failure-detection period irrespective of whether the TCP connection dies silently within or outwith the connection pool.
Edit:
KeepAliveTimeandKeepAliveCountare not valid connection string parameters in mssql-jdbc, I misread the documentation. Instead, mssql-jdbc setsKeepAliveTime/TCP_KEEPIDLEto 30 seconds andKeepAliveIntervalto 1 second (link). So does SqlClient (link) I also added a second point: while these should prevent the pool from handing out dead sockets,TCP_USER_TIMEOUT/TCP_MAXRTmay be helpful when a connection has been issued by the pool and is being used by the client. Finally: the documentation I looked at for KeepAliveCount for Windows was wrong. Newer versions of Windows do support this after all./triage
test placeholder - to be removed
Generated by SqlClient Issue Auto-Triage for #1530 · copilot · auto · 110.1 AIC · ⌖ 13.2 AIC · ⊞ 11.9K · ◷
The timing can help separate the two cases. If
Open()succeeds quickly, SQL shows no corresponding request, and failure occurs at roughlyCommandTimeout(often plus the teardown/attention interval), that is consistent with a pooled connection whose outbound flow was silently dropped—not a slow server query.For affected workloads, correlate connection checkout/idle duration with failures and compare against any NAT/firewall idle timeout. The new pool's
Connection Idle Timeoutoption (noted above for 7.1) is worth testing in a controlled rollout: it trades some cold opens after idle periods for avoiding reuse of long-idle sockets. Keep retries targeted to transient failures and ensure commands are idempotent before retrying.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsBacklog
Describe the bug
On occasions we will see the following error
However, SQL Server shows no long running queries and is not using a lot of it's resources during the periods where this happens.
It looks to be more of an intermittent connection issue but we're unable to find any sort of root cause.
To reproduce
We're not sure of the reproduction steps. I've been unable to reproduce this myself by simulating load. From what we can tell this is more likely to happen when the pod is busy (not through just HTTP, but handling events from an external source) but equally it can happen randomly when nothing is really happening on the pod which has caused us quite a substantial amount of confusion.
Expected behavior
Either more information on what the cause might be, or some solution to the issue. I realise the driver might not actually know the issue and it may really be a timeout to it's point of view. We're not entirely sure where the problem lies yet, which is the biggest issue.
Further technical details
Microsoft.Data.SqlClient version: 3.0.1
.NET target: Core 3.1
SQL Server version: Microsoft SQL Azure (RTM) - 12.0.2000.8
Operating system: Docker Container - mcr.microsoft.com/dotnet/aspnet:3.1
Additional context
TimeoutEventfrom the metrics that are collected from the pool. On occasions when we do get them, theerror_statewill be different.145. We don't know what this means can find no information on what these relate to. I've raised a ticket with the Azure Docs team to look at this. I'll add more onto this when they happen as we've not been keeping track of the error_state codes as we're not sure if they're even relevant.ReadSniSyncOverAsync