Skip to content

Security: Papyrine/Scry

Security

docs/security.md

Security model

Scry lets a client compose queries. The design assumption is that the client is hostile: the generated code, the LINQ, and the serialized request are all attacker-controlled. Every guarantee is enforced on the server, at runtime, against the real model assembly.

Threat model

Assumed:

  • An attacker can craft arbitrary JSON and send it to the query endpoint, in a POST body or encoded into a GET URL. Both reach the same handler and are validated identically; neither form is trusted more than the other, and a URL that does not decode into a request this server can parse is rejected before anything is bound.
  • An attacker can read the generated client code, and can see any schema the explorer exposes.
  • An attacker will try to name types and properties that were never generated for them.

Sensitivity is a separate axis from access. [Sensitive] does not decide whether a member may be read — the allow-list and row policies do that — it decides how the value may travel and whether the answer may be kept. A caller allowed to read a member is allowed to read it either way; what changes is that the value never reaches an access log and the response is never stored.

Not assumed:

  • That the client-side type system constrains anything. It is a developer-experience feature, not a control.

Layers

1. Default-deny allow-list

A type is invisible unless it carries [Queryable], [QueryableView], [QueryablePoco], or [QueryableComplex]. A property is invisible if it carries [QueryIgnore], has no public instance getter, or is not a scalar, a navigation to another opted-in type, or an opted-in collection.

Adding an entity to the DbContext does not expose it. Adding a property to an exposed entity does expose it — that is the one direction where the default is open, and it is why the surface should be reviewed alongside model changes.

Collection navigations are the exception to that: they are invisible even on an exposed type until the member itself carries [QueryableCollection], so adding one to a model never widens a surface by accident. An exposed collection is aggregable, not projectable — a client can ask Any, All, Count, Sum, Average, Min, or Max about it, evaluated as a correlated subquery, but can never enumerate its rows. Every answer is a scalar, so no request can return an unbounded nested collection and the page bounds are unchanged. Inside the subquery the element type's own allow-list applies, a [QueryIgnore]d member stays hidden, and a subquery may not appear inside another subquery or inside a membership test against another source.

A collection whose element type carries a row policy is refused at startup. A policy filters a source; a subquery has none, so aggregating a policied collection would count exactly the rows the policy exists to hide. A reference navigation into a policied type is not refused but rewritten: it is read through that type's policy, so a hidden row reads as null. A [QueryableComplex] type carrying a policy is refused for the same reason and at the same point: it has no source of its own for one to filter, so the policy could never run.

A collection of values — an EF primitive collection, typically a JSON column — is exposed by the same opt-in and answers the same aggregates, with two differences that follow from an element being a value rather than a row. Its element is read with the wire's element node, which the validator accepts only where the row being read is a scalar, so it can never be used to name a whole entity; and it cannot be flattened, since the rows a flatten would produce have no members for the operators after it to name.

Inheritance is not transitive either. An opt-in attribute is read off the type it is written on and never inherited, so a subclass of an exposed type is invisible until it opts in itself. Adding one to a model exposes neither its rows nor the members it declares, and no query can narrow to it: OfType names the target as a wire string, which is resolved through the allow-list and then checked to actually derive from the type being queried — so it can only ever narrow, never widen back to a base or across to an unrelated source. A row policy is the deliberate exception: it is inherited, so a subclass cannot shed the one its base carries, and where both carry one, both apply. That inheritance runs downwards only — a policy on a derived type does not filter the base source, which returns and counts those rows as instances of the base (exposing the base's members, never the derived type's). Restricting a hierarchy however it is reached means attaching the policy to the base.

Complex types and JSON columns follow the same default-deny rule. A complex/value type — including one mapped into a JSON column — is invisible until it carries [QueryableComplex], and even then only its allow-listed scalar leaves are reachable ([QueryIgnore] still hides members). A JSON array — a collection of such a type, or a collection of values — needs the member's own [QueryableCollection] on top of that, and is aggregable exactly like any other collection. Exposing a complex type does not transitively expose anything: a nested type it references is reachable only if it too is opted in, so a JSON column cannot smuggle in an entity or an unlisted field, and a [QueryIgnore]d member stays unreadable inside an array even though EF still writes it there. Traversal into a complex member counts against MaxNavigationDepth exactly like a navigation, bounding how deeply a client can descend into nested JSON. How EF stores the type (JSON or columns) never changes what is reachable — the allow-list is built from the CLR/annotation surface, not the storage mapping.

2. A closed AST

The wire format has no node for an arbitrary method call, no node for raw SQL, and no node for a type name. The full vocabulary is:

[JsonPolymorphic(TypeDiscriminatorPropertyName = "$type")]
[JsonDerivedType(typeof(WhereOp), "where")]
[JsonDerivedType(typeof(OrderByOp), "orderBy")]
[JsonDerivedType(typeof(ThenByOp), "thenBy")]
[JsonDerivedType(typeof(SkipOp), "skip")]
[JsonDerivedType(typeof(TakeOp), "take")]
[JsonDerivedType(typeof(SelectOp), "select")]
[JsonDerivedType(typeof(SelectManyOp), "selectMany")]
[JsonDerivedType(typeof(OfTypeOp), "ofType")]
[JsonDerivedType(typeof(GroupByOp), "groupBy")]
[JsonDerivedType(typeof(DistinctOp), "distinct")]
[JsonDerivedType(typeof(ReverseOp), "reverse")]
[JsonDerivedType(typeof(JoinOp), "join")]
[JsonDerivedType(typeof(SetOp), "set")]
[JsonDerivedType(typeof(CountOp), "count")]
[JsonDerivedType(typeof(LongCountOp), "longCount")]
[JsonDerivedType(typeof(AnyOp), "any")]
[JsonDerivedType(typeof(AllOp), "all")]
[JsonDerivedType(typeof(FirstOp), "first")]
[JsonDerivedType(typeof(SingleOp), "single")]
[JsonDerivedType(typeof(LastOp), "last")]
[JsonDerivedType(typeof(AggregateOp), "aggregate")]
[JsonDerivedType(typeof(PageOp), "page")]
public closed record QueryOp;

snippet source | anchor

[JsonPolymorphic(TypeDiscriminatorPropertyName = "$type")]
[JsonDerivedType(typeof(MemberNode), "member")]
[JsonDerivedType(typeof(ElementNode), "element")]
[JsonDerivedType(typeof(ConstNode), "const")]
[JsonDerivedType(typeof(BinaryNode), "binary")]
[JsonDerivedType(typeof(UnaryNode), "unary")]
[JsonDerivedType(typeof(CallNode), "call")]
[JsonDerivedType(typeof(ConditionalNode), "conditional")]
[JsonDerivedType(typeof(SubqueryNode), "subquery")]
[JsonDerivedType(typeof(CollateNode), "collate")]
[JsonDerivedType(typeof(InSourceNode), "inSource")]
[JsonDerivedType(typeof(AggregateNode), "aggregate")]
[JsonDerivedType(typeof(GroupKeyNode), "groupKey")]
[JsonDerivedType(typeof(CompositeKeyNode), "compositeKey")]
public closed record Node;

snippet source | anchor

/// <summary>The closed set of functions a client may call on a value. No free-form method names.</summary>
public enum KnownFunction
{
    StringContains,
    StringStartsWith,
    StringEndsWith,
    StringToLower,
    StringToUpper,
    StringIsNullOrEmpty,
    StringIsNullOrWhiteSpace,
    StringLength,
    StringTrim,
    StringTrimStart,
    StringTrimEnd,
    StringSubstring,
    StringIndexOf,
    StringReplace,

    /// <summary>
    /// The first and last character of a string, as <c>FirstOrDefault</c> and <c>LastOrDefault</c>
    /// spell them — a substring of one, taken at either end. The indexer that looks like it means the
    /// same is not carried: no provider translates it, and one that reads past the end of the text
    /// would fault where these answer with the default.
    /// </summary>
    StringFirst,
    StringLast,
    DateYear,
    DateMonth,
    DateDay,
    DateHour,
    DateMinute,
    DateSecond,
    DateMillisecond,
    DateDayOfYear,

    /// <summary>
    /// The sub-millisecond parts, each within the one above it: 0-999 microseconds of the
    /// millisecond, 0-999 nanoseconds of the microsecond. SQL Server's DATEPART counts them from the
    /// whole second, so the server takes the remainder, exactly as EF does.
    /// </summary>
    DateMicrosecond,
    DateNanosecond,

    /// <summary>The count of days since 0001-01-01 (<c>DateOnly.DayNumber</c>).</summary>
    DateDayNumber,

    /// <summary>
    /// The day of the week, numbered as <see cref="System.DayOfWeek"/> does — 0 for Sunday. The server
    /// owns how that is expressed in SQL, since the obvious formulation is not deterministic.
    /// </summary>
    DateDayOfWeek,
    DateDate,

    /// <summary>
    /// The time of day a date carries, as the <see cref="System.TimeSpan"/> since midnight. The
    /// counterpart of <see cref="DateDate"/>, which drops the same part instead of keeping it.
    /// </summary>
    DateTimeOfDay,

    /// <summary>
    /// The parts of an elapsed time, each within the unit above it — the hours of the day, the
    /// minutes of the hour, and so on down. Whole totals (<c>TotalHours</c> and its siblings) are a
    /// division rather than a part and no provider translates them, so they are not carried.
    /// </summary>
    TimeSpanHours,
    TimeSpanMinutes,
    TimeSpanSeconds,
    TimeSpanMilliseconds,
    TimeSpanMicroseconds,
    TimeSpanNanoseconds,

    /// <summary>
    /// Reading one temporal type as another: the date or the time half of a timestamp, a time read as
    /// an elapsed time, and a date and a time composed back into one. Each is a conversion the
    /// database performs, so the answer does not depend on the client's calendar or its clock.
    /// </summary>
    DateOnlyFromDateTime,
    TimeOnlyFromDateTime,
    TimeOnlyFromTimeSpan,
    DateTimeFromDateAndTime,

    /// <summary>
    /// Unix time, counted from 1970-01-01 UTC (<c>DateTimeOffset.ToUnixTimeSeconds</c>). The
    /// <c>DateTime</c> / <c>UtcDateTime</c> / <c>LocalDateTime</c> readings of an offset are not
    /// carried alongside them: the provider has a translation only for a column whose store type is
    /// <c>datetimeoffset</c> and refuses the expression otherwise, and the local reading would go
    /// through <c>CURRENT_TIMEZONE_ID()</c> — the server's own zone — even where it does translate.
    /// </summary>
    UnixSecondsFromOffset,
    UnixMillisecondsFromOffset,

    DateAddYears,
    DateAddMonths,
    DateAddDays,
    DateAddHours,
    DateAddMinutes,
    DateAddSeconds,
    DateAddMilliseconds,
    /// <summary>
    /// Joins the target and the argument into one string, converting either if it is not one already.
    /// C# writes this as <c>+</c>, but the operator alone does not say it: an Add of a string and a
    /// number is a concatenation, while an Add of two numbers is arithmetic, and only the client can
    /// tell which was written.
    /// </summary>
    StringConcat,

    /// <summary>
    /// The target's value as text — <c>ToString()</c> with no arguments. The formatted overload is not
    /// part of the set: no provider translates it, and the SQL function that would express it reads
    /// the server's language, so the same row would format differently per connection.
    /// </summary>
    StringFrom,

    MathAbs,
    MathCeiling,
    MathFloor,
    MathRound,
    MathTruncate,
    /// <summary>
    /// The sign of the target: -1, 0, or 1. The server composes it from comparisons rather than from
    /// SQL's own function, whose result takes the argument's type and so cannot be read back as the
    /// <see cref="int"/> this returns.
    /// </summary>
    MathSign,

    MathSqrt,
    MathPow,
    MathExp,

    /// <summary>Natural logarithm, or — with one argument — the logarithm to that base.</summary>
    MathLog,
    MathLog10,
    MathSin,
    MathCos,
    MathTan,
    MathAsin,
    MathAcos,
    MathAtan,

    /// <summary>The angle whose tangent is the target over the argument (<c>Math.Atan2(y, x)</c>).</summary>
    MathAtan2,

    /// <summary>
    /// The greater / lesser of the target and the argument (<c>Math.Max</c> / <c>Math.Min</c>). The
    /// server composes each from a comparison rather than using SQL's GREATEST and LEAST, which exist
    /// only from SQL Server 2022; a null operand keeps the answer null.
    /// </summary>
    MathMax,
    MathMin,

    /// <summary>
    /// Degrees to radians and back (<c>double.DegreesToRadians</c> / <c>RadiansToDegrees</c> —
    /// statics on the floating types rather than on <c>Math</c>). Defined over double alone, so the
    /// target is widened to reach them.
    /// </summary>
    MathDegreesToRadians,
    MathRadiansToDegrees,

    /// <summary>
    /// Membership of a client-supplied set (SQL <c>IN</c>). The target is the value being tested and
    /// every argument is a <see cref="ConstNode"/>; the server caps the number of values.
    /// </summary>
    In,

    /// <summary>
    /// Whether the target — a [Flags] enum member — carries the argument's bits
    /// (<c>Enum.HasFlag</c>). A combined flag travels by name exactly as <c>Enum.ToString</c> spells
    /// it: <c>"Parking, Gym"</c>.
    /// </summary>
    EnumHasFlag,

    /// <summary>
    /// Reads text as a value — <c>int.Parse</c> / <c>Convert.ToInt32</c> and their siblings; the
    /// inverse of <see cref="StringFrom"/> — or, over a number, widens it to the named type, which
    /// is how a cast such as <c>(double)member</c> travels. Only a widening conversion is carried: SQL's
    /// numeric-to-numeric conversions truncate where the CLR's round, so a narrowing one would answer
    /// differently per source. Text that does not parse faults at execution, exactly as it would in
    /// memory.
    /// </summary>
    Int32From,
    Int64From,
    DecimalFrom,
    DoubleFrom,
    BooleanFrom,
    ByteFrom,
    Int16From,
    SingleFrom,

    /// <summary>
    /// Three-way comparison (<c>a.CompareTo(b)</c>, <c>string.Compare(a, b)</c>): -1, 0, or 1, or
    /// null when either operand is — a comparison against a value that is not there has no direction.
    /// Numbers, text and dates compare; text compares under the server's collation, exactly as its
    /// ordering does.
    /// </summary>
    CompareTo,

    /// <summary>
    /// Questions about a binary member's bytes, without reading them: how many there are
    /// (<c>DATALENGTH</c>), whether a byte is among them (<c>CHARINDEX</c>), and the byte at one
    /// position. An <c>[Attachment]</c> answers none of them — its value is the one thing no query
    /// reads — so these reach a plain or <c>[BinaryTransfer]</c> member only. <c>Any()</c> is absent
    /// because the provider refuses it; ask whether <see cref="BytesLength"/> is above zero, which is
    /// the same question and does translate.
    /// </summary>
    BytesLength,
    BytesContains,
    BytesElementAt
}

snippet source | anchor

Unknown discriminators fail deserialization rather than being ignored, so a request that names anything outside these sets is rejected at the JSON layer.

3. Server-side revalidation

The server rebuilds the allow-list at startup from the real model assembly, independently of whatever the client was generated against. QueryValidator then walks every incoming AST and rejects:

  • An unknown root source.
  • A property that is not allow-listed on the type reached so far.
  • Traversal through a non-navigation member (Name.Length).
  • A wire version newer than the server understands.
  • An ill-formed pipeline: ThenBy without OrderBy, an operator after a terminal, more than one GroupBy or Select, Where/OrderBy after GroupBy or Select, GroupBy without a following Select, a terminal predicate after a Select.
  • An aggregate outside a grouped Select, or a grouped projection referencing a non-key member.
  • An empty projection, or a projection leaf that is not a scalar.
  • Any resource limit overrun.

Validation runs to completion before any expression is rebound or executed. A rejected query never reaches EF Core.

The optional schema stamp on a request is not a security input. It is attacker-controlled like the rest of the wire, is never consulted while deciding whether a query is allowed, and cannot widen the allow-list or unlock a source. It is read only after a query has already been rejected, to add a "the client looks stale" note to the error message. Forging or omitting it changes nothing about what a client may query.

[Test]
public Task RejectsIgnoredProperty() =>
    AssertRejected(QueryRequest.Create(
        "Employee",
        [
            new WhereOp(new BinaryNode(
                BinaryOp.GreaterThan,
                new MemberNode(["Salary"]),
                new ConstNode("100", ClrTypeTag.Decimal)))
        ]));

snippet source | anchor

4. Typed rebinding

CLR types are introduced only from the schema, never from the wire. Member access is built by looking the name up in the allow-list and using the PropertyInfo found there, so there is no path from a wire string to a reflected member that was not already allow-listed.

Collations are the exception that proves the rule. A collation cannot be a query parameter — a provider emits it into the SQL text — so a request never carries one. A client asks only for a case sensitivity; which collation implements it is server configuration, and a server that has configured none rejects the request. That keeps the invariant below intact: no attacker-supplied string reaches SQL as anything but a parameter.

Constants are the one attacker-supplied value that reaches the query. They travel as a string plus a type tag and are parsed into the member's type at the comparison site — not into whatever type the tag claims.

A parsed value is then emitted the way a captured variable reaches a query — a member read off a holder object — which is the shape EF Core's funcletizer lifts into a query parameter. The value is bound, never written into the statement text.

That shape is deliberate and worth keeping. A bare Expression.Constant is not parameterized: EF inlines it into the SQL, escaped by the provider's type mapping. Escaping makes that safe from injection, but it makes the statement text differ per value, so every distinct value a client sends compiles and caches a plan of its own — a cheap way for a hostile client to flood the plan cache. Binding gives one plan for every value. A null is the exception and stays a literal: there is one of it, so nothing is gained, and a literal null keeps EF's IS NULL rewriting straightforward.

5. Row policies

An IReturnablePolicy<T> is applied to the source before any client operator, so client filters can only narrow an already-authorized set.

A join and a membership test against another source both resolve their second source through the same path: that source's policy is applied before the two sides meet, so a join can only narrow and never becomes a way to observe rows through a source whose policy hides them.

A reference navigation into a policied type is the third route to the same rows, and is closed the same way: the traversal is rebound to read through the target's policy, so a hidden row reads as null wherever the path appears — a projection leaf, an ordering, a key, and above all a predicate, which runs in SQL and would otherwise answer about rows a direct query could never return. Because every rooted member path is rebound through one place, this holds for all of them rather than per operator. Startup translates each such traversal once and refuses to start if a policy does not compose there.

A collection navigation of a policied type is the fourth, and is refused at startup unless the policy says how it wants to be read through: an aggregate off the owner has no source for a policy to filter, so it would count exactly the rows the policy hides. Opting into DeniedCollectionMode.Hide rewrites the collection into the same policy-filtered subquery a reference navigation uses, which makes an aggregate over it count what a direct query of the element source would have reached, and a flatten reach exactly those rows.

A flatten replaces the rows with the collection's elements, and the policy chain the query carries with the element's own — the whole chain where the element is policied and read through it, none otherwise. A narrowing after it therefore applies the derived type's policies exactly as rooting at the derived source would, rather than skipping as many of them as the root happened to carry.

See Row policies.

Reporting a denial discloses that there was one

A denied row is hidden by default, at every one of those positions, and hiding is the only answer that discloses nothing. A policy can be configured to fail the request instead, per position — a 403 carrying a fixed message that names no source, member, row, or policy.

That is a deliberate existence oracle. A caller that receives it learns that rows it may not see matched its query, which is exactly the signal hiding exists to withhold; by varying the query it can narrow down what those rows are. Enable it only where "you lack permission" is itself not sensitive — an internal tool, an auditable tenant — and never on a source whose row existence is the secret. A row another policy already hid is never reported, so raising one policy's mode cannot expose what a different one is hiding, and the outcome is audited as Denied so the disclosure is countable.

Cached decisions are server-held state

A cached row policy moves the decision off the query and remembers it, keyed by policy, scope, and row key. Three things follow. The scope key must come from the authenticated principal via context.Services and never from a request header — it selects which set of answers applies, so a caller choosing it is a caller choosing its permissions. A row that is new or has changed is decided on its first read, so nothing is served on the strength of an answer nobody made. And a permission change reaches queries only when the host says it has, so the decisions can be stale by design: the lag is bounded by how promptly InvalidateRows/InvalidateScope are called, which makes calling them part of the authorization path rather than a cache optimization.

Decisions are made over the raw source rather than through its other policies, so one caller's view is never baked into an answer the others read; the other policies still apply to the query itself.

An IAttachmentPolicy<T> is the same idea for a value no query carries. A source exposing an [Attachment] refuses to start without one — the fetch endpoint is reached by row key rather than by a composed query, so the allow-list that stands between a caller and everything else has nothing to say about it. The row is still resolved through its source's row policies, so both apply, and a refusal is indistinguishable from a row that was never there.

6. Resource limits

/// <summary>Maximum number of rows a single query may request via <c>Take</c>. Default 1000.</summary>
public int MaxPageSize { get; set; } = 1000;

/// <summary>
/// Page size applied to a paged query (<c>ToPageAsync</c>) that does not request one. Bounds an
/// otherwise-unbounded page; the effective size is always capped by <see cref="MaxPageSize"/>. Default 100.
/// </summary>
public int DefaultPageSize { get; set; } = 100;

/// <summary>Maximum navigation-path length allowed in a member expression. Default 4.</summary>
public int MaxNavigationDepth { get; set; } = 4;

/// <summary>
/// Maximum number of operators in a query pipeline. The pipeline a join's inner side or a set
/// operand carries is bounded by the same number. Default 32.
/// </summary>
public int MaxPipelineLength { get; set; } = 32;

/// <summary>Maximum expression nesting depth in a predicate. Default 32.</summary>
public int MaxExpressionDepth { get; set; } = 32;

/// <summary>
/// Maximum number of expression nodes one request may carry, counted across every predicate,
/// key, selector, and projection in it — the pipeline's own, a join's or set operand's, and
/// what a subquery or aggregate reads. Default 4096.
/// </summary>
/// <remarks>
/// Depth bounds how deeply an expression nests and width how many members a projection names;
/// neither bounds a flat chain of thousands of comparisons in one predicate, which is one
/// statement the provider has to compile. The count is shape exactly as the pipeline length is.
/// </remarks>
public int MaxExpressionNodes { get; set; } = 4096;

/// <summary>
/// Maximum number of correlated subqueries one request may carry: every question about a
/// collection and every membership test against another source, wherever it appears. Default 64.
/// </summary>
/// <remarks>
/// Each is a query the database runs per row. Nesting one inside another is refused outright;
/// this bounds how many may sit side by side.
/// </remarks>
public int MaxCorrelatedSubqueries { get; set; } = 64;

/// <summary>
/// Maximum number of members a projection may name, nested members included, and the same for
/// the members a join projects. Default 256.
/// </summary>
/// <remarks>
/// Every member is an expression the provider compiles and a column the query returns, so the
/// width of a projection is work a request asks for, exactly as the length of its pipeline is. A
/// query writing no <c>Select</c> is unaffected: its projection is the model's own members.
/// </remarks>
public int MaxProjectionMembers { get; set; } = 256;

/// <summary>
/// Maximum number of values a client may supply to a set-membership test (<c>Contains</c>, which
/// becomes a SQL <c>IN</c>). Default 1000.
/// </summary>
public int MaxInValues { get; set; } = 1000;

/// <summary>
/// Maximum number of queries one batch request may carry. Default 20.
/// </summary>
/// <remarks>
/// A batch is a single request that costs more than one query, so this is the bound that keeps it
/// from being an amplifier: every other limit here is per query and would otherwise apply to an
/// arbitrary number of them. A batch over the limit is rejected whole, before any entry runs. The
/// other such request is a live query, which has bounds of its own —
/// <see cref="MaxSubscriptions"/> and the options beside it.
/// </remarks>
public int MaxBatchSize { get; set; } = 20;

/// <summary>
/// Maximum number of rows a streamed query may return, or null — the default — for no limit.
/// </summary>
/// <remarks>
/// Null matches <c>ToListAsync</c>, which has never been bounded either: <see cref="MaxPageSize"/>
/// caps <c>Take</c> and a page, not an unbounded enumeration. Nor is streaming the safer of the two
/// server-side any longer — a list that outgrows <see cref="ResponseSpillThreshold"/> is written out
/// as it is read, so neither holds its rows. What both hold is a connection and a response open for
/// as long as the client reads, which is the reason to offer a bound at all. A stream that
/// reaches the limit ends with an error marker rather than a short result, so a client cannot
/// mistake truncation for the end of the data.
/// </remarks>
public int? MaxStreamRows { get; set; }

/// <summary>
/// The longest encoded query this deployment wants asked as a URL. Default 4096; zero maps no GET
/// route at all, so every query travels as a body.
/// </summary>
/// <remarks>
/// <para>
/// Unlike the limits above this one rejects nothing — it is advertised rather than enforced,
/// because the ceiling it describes is not this server's. What actually truncates or refuses a long
/// URL is whichever hop is strictest: 8 KB on a whole request line is the common default for a
/// server or a proxy, and the number here is the budget a client is asked to stay inside of so it
/// never finds out where the real edge is. A request that arrives is answered whatever its length.
/// </para>
/// <para>
/// It is a deployment setting rather than something the model declares, since the ingress in front
/// of a server is a property of where it runs — one model can be hosted behind two of them.
/// Clients learn it from <see cref="WireFormat.UrlLimitHeader"/>, carried on every response.
/// </para>
/// <para>
/// Zero is the exception, and is enforced: it says a query may never appear in a URL here, which is
/// a statement about this deployment rather than a guess about a length. <c>MapScry</c> honours it
/// by not mapping the GET route, so routing answers such a request with a 405 naming POST and Scry
/// never sees it. Setting it means giving up conditional requests — see /docs/caching.md.
/// </para>
/// </remarks>
public int QueryUrlLimit { get; set; } = QueryUrl.MaxLength;

/// <summary>
/// Reports a query that used at least this fraction of a limit — <c>0.8</c> for eight tenths —
/// to every registered <see cref="IScryAuditor"/>, as
/// <see cref="ScryAuditEntry.ApproachedLimits"/>. Rejects nothing. Null, the default, reports
/// nothing.
/// </summary>
/// <remarks>
/// <para>
/// A limit can only be tightened once it is known how close real traffic runs to it, and a limit
/// that does nothing but reject never says: the queries that stayed inside it are exactly the
/// ones it leaves no trace of. Set this, watch for a while, then tighten on what came back.
/// </para>
/// <para>
/// It covers <see cref="MaxPipelineLength" />, <see cref="MaxExpressionNodes" />,
/// <see cref="MaxCorrelatedSubqueries" />, <see cref="MaxPageSize" />,
/// <see cref="MaxNavigationDepth" />, <see cref="MaxProjectionMembers" /> and
/// <see cref="MaxInValues" />. Two are left out. <see cref="MaxBatchSize" />, because a batch that
/// stays inside it is audited per entry rather than as a batch, so there is no entry of its own to
/// report it on. And <see cref="MaxExpressionDepth" />, because what the validator compares is how
/// many times it recursed rather than how deeply the request nests — a number this could only
/// mirror by repeating the shape of that walk, and would then misreport the day the two drifted.
/// </para>
/// <para>
/// It costs one extra walk of the request, paid only where it is set, only once an auditor is
/// registered to read the result, and only for a query that was not rejected — a refused request
/// is never measured, so this cannot be used to make refusing cost more than it does.
/// </para>
/// </remarks>
public double? LimitWatchFraction { get; set; }

snippet source | anchor

These bound the work a single request can ask for: how many rows, how deep a join chain, how long a pipeline — a join's inner side and a set operand each carry one of their own, held to the same length — how deeply nested an expression, and how wide a projection.

All but one are per query, which is what makes MaxBatchSize load-bearing: a batch is the only request that carries more than one query, so without it every other limit would apply to an arbitrary number of them at once. Each entry is otherwise validated, policy-filtered, and audited exactly as it would be sent alone — batching is a transport concern, and reaches nothing else on this page.

Three bounds are the host's rather than Scry's, and a deployment should know it leans on them. The size of a request body is Kestrel's MaxRequestBodySize — 30 MB by default — which the endpoints do not tighten themselves: put a RequestSizeLimit on the builder MapScry returns, and every endpoint it mapped is held to it, answered by the host with a 413 before a handler reads a byte. A query is small — a few kilobytes is a long one — so a limit far under the default costs nothing and bounds what a body can carry before MaxInValues refuses it, since an In list is deserialized whole before it is counted. JSON nesting is bounded by the reader at 64 levels, its default, before MaxExpressionDepth could be reached; a document past it is a malformed body and a 400. And the URL is bounded by QueryUrlLimit above, which the client also honours. HostLimitTests and HttpRoundTripTests.ADeeplyNestedBodyIsRejected pin the first two.

ResponseSpillThreshold is deliberately not among them. It decides when a response stops being resident, not how large one may be: crossing it rejects nothing, and a request that would produce a gigabyte still produces a gigabyte. What it changes is where those bytes sit while they are produced. MaxResponseBytes is the bound on how large one may be — unset by default, and the only limit that sees the size of a value rather than the number of rows and members a query asked for.

Tightening one of these is a guess until it is known how close accepted traffic runs to it — a limit that only rejects leaves no trace of the queries that stayed inside it. LimitWatchFraction is how to find out: it reports a query that came within a fraction of a limit to the audit trail, rejects nothing, and covers the four counted for a request as a whole.

7. Contained errors

Validation and wire failures return 400 with a specific message — the message names the rejected property or rule, which is not a disclosure beyond what the allow-list already implies. A type is named as the wire knows it: a source by its wire name, which is the whole point of [Queryable(Name = ...)], and a complex type by the generated model's name that introspection already publishes — never by a CLR name the allow-list chose to hide. A constant that fails to parse into its member's type is a validation failure too: parsing happens while the expression is rebound, after validation has passed, but the request is still rejected with a message naming the value rather than surfacing as a server fault. Everything else returns 500.

The 500 message is fixed — Query execution failed. — and stack traces, SQL, and EF Core messages are never returned to the client. The only variable part is the code.

A 500 is not always a server's fault, and a client can produce one on demand: a division by a constant of its own choosing, a conversion of text that is not a number, an element read past the end of a byte array. Each is a provider error at execution — after validation, which checks shape and not arithmetic — answered with the fixed message and recorded as a Failed outcome whose Error is the provider's text. Contained, but worth knowing where the failure metric feeds an alert: a caller can fill it at will, as it can the rejection count.

End to end

The generated client model has no Salary member, so a hostile client must forge the request by hand. The server rejects it:

[Test]
public async Task DisallowedPropertyRejectedWith400()
{
    const string json = """
        {
          "version": 1,
          "root": "Employee",
          "pipeline": [
            {
              "$type": "where",
              "predicate": {
                "$type": "binary",
                "op": "GreaterThan",
                "left": { "$type": "member", "path": "Salary" },
                "right": { "$type": "const", "value": "100", "tag": "Decimal" }
              }
            }
          ]
        }
        """;

    using var content = new StringContent(json, Encoding.UTF8, "application/json");
    using var response = await http.PostAsync("/api/query", content);

    await Assert.That(response.StatusCode).IsEqualTo(HttpStatusCode.BadRequest);
}

snippet source | anchor

Live queries

A live query is the same request, answered more than once. Nothing above is relaxed for it, because nothing above is skipped: every answer is the query run again through validation, the allow-list and the row policies, and recorded by the auditors. It is never a cached result, and a run is never shared between two subscriptions — sharing one across callers is the leak a row policy exists to prevent.

What is new is that the server chooses when to answer, so what has to be shown is that the choosing discloses nothing:

  • An answer is compared before it is sent. A write the caller may not see changes nothing in the rows they are allowed, so nothing goes out. When an answer arrives therefore says no more than asking again would have.
  • A heartbeat is sent on a fixed clock, from a task of its own. A run that found nothing to say cannot delay one, so a late heartbeat is not a signal that something the caller cannot see was written.
  • What reports a change names entities, never rows. Whoever can write to a backplane can cause live queries to be asked again, which costs what the throttle lets it cost, and nothing else.
  • A failure after the first answer is said as it would have been with a status: the client's own doing in full, anything else as the fixed "Query execution failed.".

Three things a live query holds for longer than a query asked once does, each bounded:

  • The authorization decision. ASP.NET Core authorizes a request once, and a live query is one request. The server ends every stream at SubscriptionLifetime, or when the authentication ticket that opened it expires if that is sooner, and the client asks again — as a new request, authorized as one.
  • Scoped services. A row policy resolved from the request's services lives as long as the subscription, and so does anything it remembered. A policy that loads a caller's grants once per scope sees a revoked grant at the next connection rather than the next run, unless the host says so: invalidating a cached policy re-runs the live queries that read that entity.
  • A policy input the query does not show. A live query runs again when something it read was written. A policy that answers by a claim, the clock, or a list it loaded in C# reads nothing the server can watch, so a change there reaches a live query at its next poll: within SubscriptionPollInterval, thirty seconds by default, and never where that is set to null.

The limits are in What it costs, and what bounds it. They are enforced by the processor rather than the endpoint, so a hub or any other transport has them too. Requests cross a hub as strings read by ScryJson, for the reason given in Hosting without the HTTP endpoint: the strictness of the wire format is in its serializer options, and a hub's own serializer has none of it.

Commands

A command is a write, and it is held to what a query is held to, plus what writing asks for:

  • Bound into the server's class. The payload arrives as JSON and is bound into the server's own command class, through the properties the server allows — unknown members refused, [CommandIgnore] properties never read, enums by name only, nested depth bounded, nulls refused where the server's class says so, DataAnnotations checked. The command's name is looked up among the server's commands; nothing the client says about its own class is trusted.
  • One answer for a target the caller may not act on. A targeted command's row is checked in one query that composes the source's row policies with the command's own: a row that is absent, one the source hides, and one the command's policy refuses are all the same 404, with the same body, so a caller probing keys learns nothing about rows it may not see.
  • A refusal is decided before anything runs. Everything a command can be refused for — malformed, unknown, denied, a target not there, one too many — is decided before the command is dispatched and before the response is committed, so each is an ordinary status and nothing ran.
  • An outcome is given again only to its caller. A command in flight, or finished within CommandRetention, is asked for again by its id. The answer goes only to the caller that sent it, by ScryOptions.Caller; anyone else, and an id this node does not hold, is told alike that it is not found.
  • Capabilities are advisory. The facade's Can* and each row's Can* member decide what a screen enables. The server decides again on every command, whatever the screen showed.
  • Off by default. MaxPendingCommands at zero maps no command route and reads every capability false. A server with commands on refuses to start while any command the model declares has nowhere to go.
  • A handler's failure is contained. A handler shows its caller only what it throws as ScryCommandException. Anything else is the fixed "Command execution failed.", and the real exception goes to the audit trail, as a query's does.

A write over a SignalR hub is not guarded the way one over HTTP is: there is no JSON content type for a cross-site form to be unable to declare, and a WebSocket handshake is not subject to CORS. A hub that authenticates by cookie keeps that cookie at SameSite=Lax or Strict, or checks the handshake's Origin; one that authenticates by bearer token is not exposed, since a browser never attaches one on another site's behalf. A hub call also has no request of its own, so a policy that reads the caller reads it from a scoped service a hub filter fills rather than from IHttpContextAccessor.

AI agents over MCP

An agent reached through MCP is one more hostile client, and it is held to exactly what any other is. It is not trusted any further because it is a model acting for a user, since what it sends can be steered by anything it has read. Every tool is a thin transport over the same ScryProcessor:

  • Off by default. ScryOptions.Mcp at Off maps no route. Read serves the schema and queries, and nothing that writes is registered, or described in the schema it is given. ReadWrite adds commands, and refuses to start where commands are off.
  • The wire format, read strictly. A query arrives as the JSON AST and is read by ScryJson, exactly as a generated client's is. Unknown operators, nodes and members are refused, and the result then goes through the validator and the allow-list.
  • No code is run. The explorer compiles C# in the browser; the MCP transport accepts no C# at all. Compiling an agent's snippet on the server would run attacker-chosen code in-process, ahead of every check the processor makes.
  • The caller is the request's. The server is stateless, so each tool call is an HTTP request of its own, answered in its own scope. Authentication, RequireAuthorization on MapScryMcp, row and command policies, and ScryOptions.Caller all see the same user they would over MapScry.
  • Rejections are said to the agent. A query refused for its own shape names what was wrong, so the agent can correct itself. Any other failure is the fixed text, as over HTTP.

Put authorization on the endpoint, and prefer a bearer token to a cookie, as MCP's own authorization does. An agent asks what its user asks, and a row policy that scopes by the authenticated principal scopes the agent too.

What Scry does not do

Authentication and authorization. Scry has no notion of a user. Put it on the endpoint:

app.MapScry("/api/query")
    .RequireAuthorization("Reader");

Rate limiting and cost control. The limits bound the shape of a query, not its cost. An allow-listed query over a large unindexed table is still expensive, and MaxPageSize caps an explicit Take rather than implicitly paging an unbounded query. Apply ASP.NET Core rate limiting, a command timeout, and the usual database-side controls. A live query is counted by rate limiting as the one request it is, whatever it goes on to run: the most it can ask of the database is bounded by MaxSubscriptions and SubscriptionThrottle rather than by a limiter, and is driven by other callers' writes.

Bound how long a slow reader can hold a connection. A response past ResponseSpillThreshold is written as it is read, so it holds a connection and its database read open for as long as the client takes to read it. That exposure is not new — …/stream has always had it, and MapScry maps every endpoint together precisely so the surface is uniform rather than one endpoint being protected while its neighbours are not — but it now reaches ToListAsync as well, which MaxStreamRows does not bound. Set the threshold to zero to hold responses whole as they once were, at the cost of an unbounded result being resident. Set MaxResponseBytes to bound how much any one response can carry, which the threshold never does. The improvement in the same change is that such a result is no longer resident twice, as rows and as serialized bytes.

Retry, order, or deduplicate commands. A refused command is the caller's to send again, a command whose outcome is Unknown may have run, and nothing orders two commands carried over a bus. A client that must not apply a command twice says so in the command — an idempotency key its handler checks — rather than relying on the transport.

Column-level authorization per user. [QueryIgnore] is static: a column is exposed or it is not. There is no per-caller column masking. Expose a view containing only the permitted columns instead.

Trusting a request header. A row policy can read the call's headers off ScryPolicyContext.RequestHeaders, and the client can attach them per query. Every one of them is chosen by the client and therefore attacker-controlled — a policy that scopes rows by X-Tenant scopes nothing, because an attacker sends a different X-Tenant. They are hint data: correlation ids, trace ids, a client build. Identity and tenancy come from the authenticated principal, resolved through context.Services.

Auditing, by default. Nothing is recorded until something subscribes. The hooks exist — every query is reported to any registered IScryAuditor with its full request AST and outcome, alongside traces and metrics — but turning them on, and alerting on rejections, is deployment work.

CORS, CSRF, TLS. Ordinary ASP.NET Core concerns, mostly unchanged by Scry. The one thing the endpoints do themselves is refuse a body that is not application/json (a 415): an HTML form can navigate a browser to a POST endpoint with a text/plain field shaped as JSON, but it cannot set that header, so requiring it keeps a cross-site page from executing a query — or fetching an attachment as a document — as whoever the browser sent. An anti-forgery token, where the host wants one, goes on top.

Cache or range-serve an attachment. Every fetch is authorized afresh, and there is no ETag, Cache-Control, or Range support — a cached attachment is one the policy no longer sees. Add caching in middleware only where that trade is acceptable.

Widen anything for binary transfer. [BinaryTransfer] changes how an already-allow-listed byte[] value is encoded in the response — a raw multipart part instead of base64 — and nothing about what a request may ask for: it adds no request-side input at all, and validation, policies, and limits are untouched. The response side is server-to-client and outside the hostile-client model; the client still bounds what it reads (the multipart reader, from the HttpMultipart package, keeps its header count/length limits), so a compromised or misbehaving server cannot make it buffer unbounded headers.

Review checklist

  • Every [Queryable] type is intended to be client-readable, and its exposed properties reviewed.
  • Sensitive columns carry [QueryIgnore] — and any newly added ones too.
  • Multi-tenant sources have a row policy.
  • A policy meant to cover a hierarchy is attached to the base, not only to a subclass — it does not filter upwards.
  • Every [Attachment] has a policy that authorizes the caller, not merely one that returns true.
  • No row policy scopes rows by a request header — those are client-chosen. Nor does a cached policy's ScopeKey, which selects a whole set of decisions.
  • Every position set to DeniedRowMode.Error is one where revealing that hidden rows exist is acceptable — hiding is the non-disclosing default.
  • Every [QueryableCollection] opted into DeniedCollectionMode.Hide is one whose aggregates are meant to be answerable at all.
  • Every grant change that a cached row policy would decide differently calls InvalidateRows or InvalidateScope — nothing else can know.
  • Where conditional requests are on, whatever a cached policy's decisions depend on is in CacheScope — invalidating the policy does not move an ETag, so a caller holding one is answered 304 with the rows it no longer has access to.
  • The query endpoint requires authentication/authorization.
  • MaxPageSize matches what the UI actually needs.
  • The explorer is either unmapped or behind a real guard in production.
  • If the explorer is exposed to anyone in production, its SQL preview is left off — the SQL discloses real table and column names and the shape of every row policy.
  • Rate limiting and a database command timeout are configured.
  • Where live queries or commands are on, Caller reads the authenticated principal, never a header — a caller that names itself is bounded by nothing, and could ask for another's command outcome by its id.
  • Commands are on (MaxPendingCommands) only on the servers meant to accept writes.
  • Every targeted command has a policy that decides its rows for the caller, not merely one that returns true.
  • A SignalR hub that carries commands and authenticates by cookie keeps that cookie at SameSite=Lax or Strict, or checks the handshake's Origin.
  • Where a row policy answers by something the query does not read — a claim, the clock, a list loaded in C# — SubscriptionPollInterval is as short as a revoked permission may be allowed to last.

There aren't any published security advisories