Skip to content
Search lessons, topics, tests…
Esc

    ↑ ↓ moveEnter openEsc close

    Module 1 · Memory and ownership in modern C# · Lesson 1 of 2

    Allocation-Aware Parsing with ReadOnlySpan<char> in C#

    Watch

    Learning outcome

    By the end of this lesson, you should be able to explain why ReadOnlySpan<char> is a strong fit for parsers, tokenizers, and protocol readers that want to avoid extra string allocations. You should also be able to describe what slicing does, when a span is just a view into existing memory, and why the compiler blocks some seemingly convenient patterns that would let a span outlive its source.

    Intuition

    Taking a non-empty proper substring normally creates a new string; whole-string and empty-string cases can reuse existing instances. That is perfectly fine in many apps, but in hot parsing paths it can create avoidable GC pressure. A ReadOnlySpan<char> is different: it is a lightweight view over a contiguous region of memory. The view contains a reference and a length; it does not own the data.

    That is why span-based parsing often looks like this: start with the original text, find delimiters, then slice the span into smaller spans for each field. Those slices still point at the same backing memory, so you can inspect pieces of the input without copying them.

    The important mental model is:

    • string is an immutable object; multiple callers can share it.
    • ReadOnlySpan<char> is a temporary window onto data.
    • slicing a span narrows the window; it does not duplicate the text.

    Deep dive

    ReadOnlySpan<char> is a ref struct, so the compiler restricts where the span value may escape. Its backing data can still live on the managed heap. That restriction is not an arbitrary annoyance; it is what keeps the type safe when it can point to stack memory, native memory, or an array. If a span could be stored in a class field, boxed, captured by a lambda, or returned beyond a safe backing lifetime, the program could end up holding a view to memory that is already gone.

    For parsing, this gives you a useful pattern:

    1. Accept ReadOnlySpan<char> as input.
    2. Locate separators with index-based scanning.
    3. Use Slice(start, length) or Slice(start) to isolate tokens.
    4. Convert only the pieces you truly need as owned values, such as int.Parse(...) or ToString() on a final token.

    A span slice is cheap because it only adjusts the window boundaries. It is not a substring copy. That makes it ideal when you need to validate or route data before materializing anything.

    Failure modes

    The most common mistake is assuming a span can behave like a regular object. It cannot. Because it is a ref struct, these patterns are blocked:

    • storing it in a class field,
    • capturing it in a lambda,
    • preserving it across an await boundary,
    • returning a span over this method’s own stackalloc local.

    Another trap is converting too early. If you call ToString() on every token, you often lose the allocation benefit. A good parser delays conversion until the point where an owned string or typed value is actually required.

    Also remember that slicing does not validate semantics; it only validates bounds. Slice(0, 4) on a five-character input is legal even if the first four characters are meaningless to your protocol.

    Interview drill

    Try answering these out loud:

    1. Why does ReadOnlySpan<char> reduce allocations compared with Substring?
    2. What does Slice(...) return, and what does it not do?
    3. Why is ReadOnlySpan<char> a ref struct instead of a normal struct?
    4. Give one example of a span lifetime error the compiler prevents.
    5. When would you deliberately convert a token span to string anyway?

    A strong answer should mention ownership, lifetime, and the difference between a view and a copy.

    Revision checklist

    • ReadOnlySpan<char> is a non-owning view over contiguous memory.
    • Slicing adjusts start/length metadata rather than copying characters.
    • Span-based parsing is useful when you want to delay allocation.
    • ref struct restrictions exist to prevent unsafe escaping of stack-bound or temporary memory.
    • Convert to owned strings only when you need persistence or API compatibility.

    Worked parser

    The example below tokenizes a simple comma-separated header and skips empty fields without creating intermediate substrings. It is not a complete CSV parser and does not reject missing fields or handle quoted commas. It also demonstrates a safe lifetime boundary: the parser consumes spans immediately and returns owned results only at the end.

    Code walkthrough

    The parser accepts ReadOnlySpan<char> instead of string. That means callers can pass an existing string with AsSpan(), but the parser itself is free to work with the text as a view.

    Each loop does three things:

    • finds the next delimiter with IndexOf(','),
    • slices out the current token,
    • trims and converts only the final token form that the result list needs.

    Notice that Trim returns another span. It does not allocate. The parser still allocates the result list, its backing storage and owned token strings. This sample makes no exact allocation-count claim. In a real parser, you could keep the tokens as spans for even longer if the downstream API also accepted spans.

    The compiler rejects escaping stack-backed spans; it does not track pool returns or unmanaged-memory release. ReadOnlySpan<T> prevents writes through that view, not changes through other aliases. The array example below distinguishes a live view from a copied string. Since C# 13, async methods can contain span work in permitted synchronous regions, but cannot preserve a span across await. Memory<T> can cross that boundary; the buffer must still remain valid.

    Executable code examples

    Span-based token parsing without substring copies

    Program.cs

    C#
    using System;
    using System.Collections.Generic;
    
    Show("basic", "id, name, status");
    Show("empty-fields", ", id, , name, ");
    Show("empty", "");
    Show("whitespace", " \t,\r\n ");
    Show("single", " status ");
    Show("unicode-space", "\u2003id\u2003,name");
    Show("quoted-comma", "\"a,b\",c");
    
    char[] source = { 'a', 'b' };
    ReadOnlySpan<char> view = source;
    string copy = view.ToString();
    source[0] = 'z';
    Console.WriteLine($"view={view.ToString()}; copy={copy}");
    
    static void Show(string name, string text)
    {
        List<string> fields = ParseFields(text.AsSpan());
        Console.WriteLine($"{name}: count={fields.Count}; [{string.Join("|", fields)}]");
    }
    
    static List<string> ParseFields(ReadOnlySpan<char> input)
    {
        var result = new List<string>();
    
        while (!input.IsEmpty)
        {
            int comma = input.IndexOf(',');
            ReadOnlySpan<char> token = comma < 0 ? input : input.Slice(0, comma);
            token = Trim(token);
    
            if (!token.IsEmpty)
            {
                result.Add(token.ToString());
            }
    
            if (comma < 0)
                break;
    
            input = input.Slice(comma + 1);
        }
    
        return result;
    }
    
    static ReadOnlySpan<char> Trim(ReadOnlySpan<char> value)
    {
        int start = 0;
        int end = value.Length - 1;
    
        while (start <= end && char.IsWhiteSpace(value[start])) start++;
        while (end >= start && char.IsWhiteSpace(value[end])) end--;
    
        return start <= end ? value.Slice(start, end - start + 1) : ReadOnlySpan<char>.Empty;
    }

    Worked interview answers

    1. Slices inspect existing characters. This implementation still creates a list and final strings; it does not demonstrate zero allocations.
    1. Slice returns a smaller view. It neither copies nor checks this header’s field rules.
    1. A ref struct enables ref-safety checks that ordinary heap-storable values cannot provide.
    1. Returning a span over a method’s stackalloc local is rejected. Returning a slice of a suitable caller-provided span can be valid.
    1. Materialize a string when the result needs independent storage or a string-only API.

    Quick reference: comparing and validating tokens

    Use left.SequenceEqual(right) to compare characters. Span == checks the start and length of a view, so it cannot answer whether separate buffers contain the same token.

    For a protocol that accepts one or more ASCII digits and an Int32 result, validate that exact grammar before numeric conversion. Allow leading zeroes if the protocol permits them; converting to an integer does not preserve their formatting. A parsing style alone is not a substitute for the protocol's character rules.

    Run this complete Program.cs with the lesson's .NET 10 / C# 14 project.

    C#
    using System;
    using System.Globalization;
    
    Show("digits", "012");
    Show("sign", "+12");
    Show("space", "12 ");
    Show("nul", "12\0");
    Show("empty", "");
    Show("overflow", "2147483648");
    Show("non-ascii", "\u0661\u0662");
    
    static void Show(string name, string text)
    {
        bool ok = TryProtocolInt(text.AsSpan(), out int value);
        Console.WriteLine($"{name}: {ok}/{value}");
    }
    
    static bool TryProtocolInt(ReadOnlySpan<char> token, out int value)
    {
        value = 0;
        if (token.IsEmpty) return false;
        foreach (char c in token)
            if (c is < '0' or > '9') return false;
        return int.TryParse(token, NumberStyles.None, CultureInfo.InvariantCulture, out value);
    }

    Expected output:

    Text
    digits: True/12
    sign: False/0
    space: False/0
    nul: False/0
    empty: False/0
    overflow: False/0
    non-ascii: False/0

    Here \0 denotes a NUL character.

    Revision note: identify the valid range, enforce the protocol grammar, then choose whether the result needs an owned value. Equality of contents, lifetime safety, and valid protocol data are separate checks.

    References: span equality, SequenceEqual, Int32.TryParse.

    Solved exercise: preserve positions?

    Input ", id, , name, " produces two values: id and name. Empty and whitespace-only fields are skipped, including the trailing field. If a protocol assigns meaning to column positions, this policy loses information. Change the parser’s contract before reusing it there; a test expecting five positions cannot be satisfied by the current implementation.

    Input "a,b",c contains a quoted comma. The sample splits it into three pieces: "a, b", c. That is a deliberate limitation, not CSV support. The final alias exercise changes the array from ab to zb: the view reflects the edit while the earlier string remains ab.

    Run and check

    Use a .NET 10 console project with C# 14. Place the listing in Program.cs. No extra packages are needed. Save the project file below as SpanParsing.csproj. Expected output:

    Text
    basic: count=3; [id|name|status]
    empty-fields: count=2; [id|name]
    empty: count=0; []
    whitespace: count=0; []
    single: count=1; [status]
    unicode-space: count=2; [id|name]
    quoted-comma: count=3; ["a|b"|c]
    view=zb; copy=ab

    Project file

    XML
    <Project Sdk="Microsoft.NET.Sdk">
      <PropertyGroup>
        <OutputType>Exe</OutputType>
        <TargetFramework>net10.0</TargetFramework>
        <LangVersion>14.0</LangVersion>
        <Nullable>enable</Nullable>
        <ImplicitUsings>disable</ImplicitUsings>
        <TreatWarningsAsErrors>true</TreatWarningsAsErrors>
      </PropertyGroup>
    </Project>

    Run dotnet run --project SpanParsing.csproj -c Release in that directory.

    Optional video

    Microsoft .NET: Stephen Toub and Scott Hanselman explore how spans represent memory. Start at 31:51 for the Span discussion. This recording is from May 2024; use this lesson's C# 14 guidance for current async/ref-struct rules. Watching is optional.

    Watch the official video on YouTube.

    Practice

    Sign in to mark lessons done and keep your place in the course.Sign in