cat ./notes/thread-local-storage-matrix.md

Thread Local Storage Matrix

Ordinum's implementation of Thread Local Storage and how the storage engine interacts with it

This journal walks through the design challenges and decisions faced with when using thread local storage for the Ordinum storage engine.

At a high level, thread local storage is quite simply what it's title suggests. A per thread process with storage local to the thread that can be called into and accessed for that thread only, independent of other thread processes. In rust this comes in the form of the thread_local!() macro

macro_rules! thread_localrust

macro_rules! thread_local {
    () => { ... };
    ($($tt:tt)+) => { ... };
}

It implements std::thread::LocalKey which is a key into the underlying storage for tls on the target platform. For example, on Linux that maybe the TLS support of the ELF ABI (..).

LocalKey uses the fastest implementation available on the target platform and is constructed with the thread_local! macros as described above.

Some interesting points when it comes to TLS and LocalKey:

  • Initialisation is done dynamically/lazily on the first call to a setter X.with(..)
  • Although TLS is a single thread primitive, it is possible for thread local state to be shared with other threads so the implementation detail specifies that only &T references may be obtained. It is therefore necessary to encapsulate tls fields with interior mutability primitives such as Cell<>, UnsafeCell<>, RefCell<> etc. if mutability is required.
  • Destructors are 'best effort' and platform specific. A number of caveats are known for where destructors are not run (..)

Here is an example of how we would initialise and interact with thread local storage in rust:

thread_local!()rust

use std::cell::{Cell, RefCell};

thread_local! {
    pub static FOO: Cell<u32> = const { Cell::new(1) };

    static BAR: RefCell<Vec<f32>> = RefCell::new(vec![1.0, 2.0]);
}

assert_eq!(FOO.get(), 1);
BAR.with_borrow(|v| assert_eq!(v[1], 2.0));

The book Rust Atomics and Locks by Mara Bos has an excellent introduction into threads and concurrency for low level systems. Highly worth a read.

Ordinum makes use of thread local storage quite heavily, not just for capturing metrics but as subsystems for optimisations and efficiencies.

Problem Statement

Ordinum has a number of sub-systems which require the use of thread local storage to speed up processes and behaviour, and to reduce strain on global structures. It also needs thread local storage to capture per process metrics and local state which extend for the lifetime of the thread. Both of these are orthogonal to each other. The former must ineract with the program with state having different lifetimes and accessors based on the program logic, for example scoped to per database instances. The latter, stretches for the length of the thread process and is purely local to that thread.

For the storage state which must interact with the program, the problem becomes, how do we effectively separate state from different instances of the program and protect cross thread interaction.

Those are the two axis which we are to focus on for this.

tls_axis

If we cast our mind back up to the TLS/LocalKey invariants, it is mentioned that we only are given &T references back from TLS and that we must carefully address the fact that other threads can (and in our case, will) touch thread local storage and more importantly may mutate thread entry state in certain cases.

Ordinum will have subsystems of varying complexity and needs which will need to utilise tls and this problem space is what we'll address further in the journal.

Naive Implementation

Where possible, it is recommended to start with the naive implementation. Although my perfectionism often overules this and forces me into an optimisation blender. For Ordinum's thread local storage, we truly started with the basic implementation and discovered along the way what needed to be changed based on the problem evolving as we introduced more complex subsystems and invariants.

We will talk to 3 subsystems (non-exhaustive), each with their own needs and each covering the different problems as described.

  1. PerfContext

- Local thread metrics

  1. BatchCache

- A cache local to the thread which stores an array of batches for grouped writes

  1. SuperVersion Cache

- A cached pointer to the superversion subsystem for snapshot reads

To begin, it is useful to briefly outline why we decide to use thread local storage, the decision as to why we would want to use it and how Ordinum can makes use of it to create efficiencies.

Why We use Thread Local Storage

We can see how thread local storage might be useful by starting with the example of the write batch. For writers writing to the database engine, at it's simplist, we can imagine a single thread/process which takes bytes and calls into the storage engine to write them.

The storage engines job is to carry out the write. Our job is to make sure the storage engine does this efficiently and the writer is not stuck or waiting a long time during this process.

The storage engine will seek to batch the writes, again thinking simply, this will take the form of allocating a contigous chunk of memory like a Vec<U8> to house the batched writes and run it through it's pipeline process

mermaidmermaid
flowchart TD W[Writer] --> D["DB::write()"] --> B["Allocate New Batch<br/>Vec&lt;u8&gt;"] --> P["Commit to Write Pipeline"]

Now, clearly looking at this you may think, 'Well, why not create a pool of batches?' and you'd be right! So the next optimisation would be to create a global pool scoped to the database instance. Of course, there'd be many different ways you could go about this, ways which are outside the scope of this journal.

We create a global pool whereby we lazily allocate batches and return them to the pool on drop.

This works fine, and for most workloads this is an already optimal solution. But, thinking about thread local storage, can we do better?

Yes! We can and we shall.

If we imagine a very basic batch pool looking somehting like:

BatchPoolrust
struct BatchPool {
  batches: Mutex<Vec<Batch>>,
}

We have a shared structure which many processes will be hitting. We have a Mutex<_> which all those processes will be trying to acquire and will be waiting for if they do not have the lock.

The Write Path is most definately a hot path and so we want to reduce contention as much as possible, where possible.

mermaidmermaid
flowchart TD Writer[Writer] Writer --> Acquire["Acquire Batch"] --> Pool[(Global Batch Pool)] Pool --> Commit["Commit to Write Pipeline"] Commit --> Return["Return Batch"] Return --> Pool

The solution is to lean on thread local storage to allow threads to avoid hitting global structures on the hot path and thereby avoid contention. We implement a caching structure in thread local storage where allocated batches are cached and can be reused by threads on writes.

The process flow can be reduced to:

  1. Thread checks it's own cache
  2. If empty, check the global pool
  3. If empty, allocate new batch

And on a thread finishing with a batch, the reverse of this is:

  1. Try to return to thread local cache
  2. If full, try to return to global
  3. If full, destroy the batch
mermaidmermaid
flowchart TD W[Writer] --> A["Acquire Batch"] A --> TLS[(Thread Local<br/>Batch Cache)] TLS -->|Hit ✓| P["Write Pipeline"] TLS -->|Miss| GP[(Global Batch Pool)] GP -->|Hit| P GP -->|Miss| N["Allocate New Batch"] N --> P

Building Thread Local Storage for Ordinum

The first introduction I had to thread-local storage (TLS) came while tackling what is probably every systems programmer's worst nightmare: memory reclamation.

I was working on a problem that required the use of Epoch based reclamation While that's outside the scope of this journal, it was certainly a rude awakening to the realities of concurrent state, synchronisation, and designing data structures that can safely outlive the threads accessing them.

For Ordinum, the first iteration of TLS looked almost exactly like the introductory example above. We had a simple registry module containing a thread context that stored any per-thread objects we wanted to cache or reuse.

This implementation served its purpose well. It allowed the design to evolve naturally while answering questions such as: Why are we using TLS here? When does it make sense to cache data on a per-thread basis? When is TLS the wrong solution?

Before long, however, a more fundamental architectural question emerged:

What happens if an application opens multiple database instances?

This was the first major design decision surrounding the TLS implementation. Should Ordinum even support multiple database instances?

There are perfectly reasonable arguments for supporting only a single database per process, and many applications will never need anything more. However, allowing multiple independent database instances provides an important level of operational isolation. Each database owns its own storage engine, write pipeline, WAL, compaction scheduler, caches, and configuration, allowing different workloads to coexist without interfering with one another.

This naturally raises another question: couldn't these simply be column families instead?

The answer is that column families solve a different problem.

Database Instances vs Column Families

A column family is a logical namespace within a database. It has its own memtable and LSM tree, but shares the database's infrastructure.

A database instance is a completely independent storage engine with its own WAL, write pipeline, background workers, caches, and metadata. (Although the benefit of database instances on the same machine is that they can still share and utilise program wide infrastructure such as thread pooling).

RequirementColumn FamilySeparate Database
Logical separation✅✅
Separate LSM tree✅✅
Separate memtables✅✅
Shared WAL✅❌
Shared write pipeline✅❌
Shared sequence numbers✅❌
Shared background compactions✅❌
Independent configuration❌✅
Failure isolation❌✅

We would ideally use a comlumn familty when the datasets belong together and should share infrastructure.

  • Users
  • Orders
  • Products

And use a database when the datasets are operationally independent.

DatasetRetentionCompressionImportance
UsersForeverZstdCritical
Cache1 hourNoneDisposable
Logs30 daysLZ4Medium

For example, if the cache database fills the disk or experiences heavy write stalls, the user database can continue operating unaffected.

But again, this is purely a matter of opinion, and only you know the needs of your workload, Ordinum just supplies the concise tools and engine to service those.

I digress. Ordinum ultimately chose to support mutliple database instances, and this brings us to the first iteration of that.

Thread Local Matrix

ThreadContext Instancesrust
thread_local! {
    static THREAD_CTX: UnsafeCell<ThreadContext> =
        UnsafeCell::new(ThreadContext::new());
}

/// Main thread local storage structure
struct ThreadContext {
  //
  instances: Mutex<HashMap<usize, DBInstanceContext>>,
  //
  // Other fields not relevant ...
}

/// Structure for each database instance stored within thread local storage
struct DBInstanceContext {
  //
  batch_cache: Vec<Batch>,
  //
  // Other tls sub sytems ...
}

Our first approach was to simply define the database instance as a seperate structure which we would store in a HashMap inside the ThreadContext this gives the benefit of being very simple and very clear in it's intention.

In the code example we have the DBInstanceContext struct which houses the batch_cache. If we were to put another sub-system in there we can see how we might be encapsulating state within the context of the db instance sort of like a registry.

ThreadContext Instancesrust
/// Structure for each database instance stored within thread local storage
struct DBInstanceContext {
  //
  batch_cache: Vec<Batch>,
  //
  superversion_cache: SVCache,
  perf_context: PerfContext,
  //...
}

For this approach we needed an ID to be able to access each database instance and retrieve it's context of sub-systems. This is simple enough, on each DB::Open() we increment a global const Atomic number DB_ID.fetch_add(1, Ordering::Release). And on each thread we lazily add to the HashMap on first access to TLS.

The key mental model to hold is that we (at this point) are storing the DBInstanceContext object as a whole, including all sub-systems within. So when we access thread local data, we must go through the context. This is ok for simple cases such as referencing/reading the data inside but for operations on sub-systems that might require different lifetimes or mutability contracts, we run into problems.

mermaidmermaid
flowchart TD DB["DB Instance<br/>db_id = 3"] DB -->|"lookup using db_id"| TC subgraph TLS["Thread-Local Storage"] direction TB TC["ThreadContext"] MAP["HashMap&lt;DbId, DBInstanceContext&gt;"] CTX["DBInstanceContext"] TC --> MAP MAP -->|"db_id = 3"| CTX subgraph SUB["Per-Database TLS Subsystems"] direction TB BC["Batch Cache"] SV["SuperVersion Cache"] OTHER["Other TLS Subsystems"] end CTX -->|"context.batch_cache"| BC CTX -->|"context.sv_cache"| SV CTX -->|"context.other"| OTHER end CTX --> ISSUE["Single context owns every subsystem<br/><br/>Different subsystems need different<br/>borrowing, lifetimes and mutability"]

The "aha!" (or perhaps more accurately, the "oh no... time to refactor") moment came while thinking about how the objects inside DBInstanceContext would actually be destroyed.

At some point the database instance is closed, meaning every thread-local subsystem stored within the context must eventually be dropped. The question quickly became: when is it actually safe to do that?

Initially it seemed reasonable that the DBInstanceContext itself should own this responsibility. However, the more I thought about it, the less practical that became. Different subsystems have completely different lifetime requirements. A thread-local batch cache, for example, has a very different notion of "safe to destroy" than a SuperVersion cache holding protected pointers.

To make destruction safe from the context itself would require introducing in-flight operation counters, additional reference counting, or other coordination mechanisms so that the context could somehow know when every subsystem had reached a quiescent state. Even then, the context would be making the flawed assumption that every subsystem follows the same shutdown semantics.

That was the realisation: DBInstanceContext was owning far too much. Rather than treating every subsystem as though it shared a common lifetime model, each subsystem should own its own lifecycle and define for itself what "safe to destroy" actually means.

Luckily, this problem had already been encountered and implemented by RocksDB who design the thread local storage structure as a matrix. (thread_local.cc)

texttext
This is the structure that is declared as "thread_local" storage.
The vector keep list of atomic pointer for all instances for "current"
thread. The vector is indexed by an Id that is unique in process and
associated with one ThreadLocalPtr instance. The Id is assigned by a
global StaticMeta singleton. So if we instantiated 3 ThreadLocalPtr
instances, each thread will have a ThreadData with a vector of size 3:
     ---------------------------------------------------
     |          | instance 1 | instance 2 | instance 3 |
     ---------------------------------------------------
     | thread 1 |    void*   |    void*   |    void*   | <- ThreadData
     ---------------------------------------------------
     | thread 2 |    void*   |    void*   |    void*   | <- ThreadData
     ---------------------------------------------------
     | thread 3 |    void*   |    void*   |    void*   | <- ThreadData
     ---------------------------------------------------

Similar to the diagram at the beginning of this journal, each thread is represented by a row in the matrix. The important difference is that the meaning of each column has changed. Previously, each column represented a database instance, with a DBInstanceContext acting as a container for every thread-local subsystem associated with that database.

Instead, each column now represents a single thread-local subsystem instance. Rather than assigning a DB_ID, each subsystem registers itself and is assigned a unique TLS_ID. This TLS_ID is generated by a global allocator and used as the index into each thread's entries vector. On first access, the vector is lazily resized and a pointer to that subsystem's thread-local state is stored directly in the corresponding slot.

The immediate benefit is that we no longer have to access a subsystem by first retrieving a DBInstanceContext. Each subsystem can be accessed directly through its own TLS_ID, allowing it to own its own initialization, lifetime, and destruction semantics independently of every other subsystem.

This may initially seem like a subtle distinction, but it becomes much more powerful as additional thread-local subsystems are introduced. Consider the BatchCache. Since each database owns a single batch cache, it is natural for each database instance to contribute one column to the matrix.

The advantage becomes much clearer with the SuperVersionCache. Unlike the batch cache, SuperVersions exist on a per-column-family basis. A single database may contain many column families, each requiring its own cached SuperVersion. Under the original DBInstanceContext design, these independent caches would all have been hidden behind a single context object. With the new model, each cache simply registers its own TLS_ID, resulting in one column per cached SuperVersion instance. The thread-local matrix does not need to understand what the subsystem represents it simply provides fast, direct access to thread-local state.

texttext
                         Database 1
                  ┌──────────────────────┐
                  │ BatchCache           │
                  │ SVCache(CF0)         │
                  │ SVCache(CF1)         │
                  │ SVCache(CF2)         │
                  └─────────┬────────────┘
                            │
                            ▼

        ------------------------------------------------------------------------------------------------------------
        |          | BatchCache(DB1) | SVCache(CF0) | SVCache(CF1) | SVCache(CF2) | BatchCache(DB2) | ... |
        ------------------------------------------------------------------------------------------------------------
        | thread 1 |      void*      |    void*     |    void*     |    void*     |      void*      |     |
        ------------------------------------------------------------------------------------------------------------
        | thread 2 |      void*      |    void*     |    void*     |    void*     |      void*      |     |
        ------------------------------------------------------------------------------------------------------------
        | thread 3 |      void*      |    void*     |    void*     |    void*     |      void*      |     |
        ------------------------------------------------------------------------------------------------------------

The next problem is 'how do we implement this?'

Implementation

We need to think about the two axes of the thread-local matrix: columns and rows. Each represents a different access path and therefore serves a different purpose.

As we've discussed already, a column represents a distinct subsystem (or vertical), independent of its neighbouring columns. A row represents a thread and contains that thread's data for each registered subsystem.

An easy way to visualise this is as a spreadsheet. To access a particular cell, we must first identify the column (the subsystem), then traverse to the required row (the thread). The intersection of the two gives us the object we wish to access.

This naturally gives us two traversal directions:

  • Across rows – iterate over every thread for a particular subsystem.
  • Across columns – iterate over every subsystem belonging to a particular thread.

Before diving into the implementation, it's worth understanding why these traversal paths are necessary.

We'll use two examples:

  1. Database shutdown
  2. Thread exit

Database Shutdown:

In our earlier BatchPool example, we described how each pool is scoped to a database instance. Multiple threads may be accessing that pool simultaneously, meaning each thread's row contains a cell for that pool's thread-local cache.

When the database shuts down, the pool itself is destroyed. Consequently, the entire column representing that subsystem must be removed to prevent threads from accessing caches belonging to a database that no longer exists.

To achieve this, the shutdown thread traverses every row and invokes the registered destructor for that column. Only the cells belonging to the shutting-down database are removed; all other columns and thread-local subsystems remain untouched.

Thread Exit:

Thread exit is where we can expect to experience the highest churn. Many threads may come and go, be spun up and destroyed, or equally be long running.

When a thread exits, we instead traverse across that thread's row. For every populated cell, the registered destructor is invoked before the entry is cleared. Once complete, the entire row can be safely reclaimed.

This operation is performed while holding the global thread registry mutex. The mutex serialises thread registration and removal, ensuring that database shutdown and thread exit cannot race while traversing the registry.


To start with, we can begin looking at what data threads will store and how it is structured. We begin with two high level objects:

  1. ThreadMetaGlobal
  2. ThreaData

..


Each row is owned by a particular thread but is discoverable by other threads through the global registry. Accessing another thread's row is an uncommon operation used for coordination tasks (e.g. cache invalidation or reclamation) and therefore requires carefully designed synchronization and ownership invariants.

To do this we use a Doubly Linked List where we start at a sentinel node and traverse registered threads

mermaidmermaid
flowchart TB COL["Column = Batch Pool"] --> ROW1["Thread A"] --> ROW2["Thread B"] --> ROW3["Thread C"] ROW1 --> CELL1["BatchCache"] ROW2 --> CELL2["BatchCache"] ROW3 --> CELL3["BatchCache"] CELL2:::target classDef target fill:#c8e6c9,stroke:#2e7d32,stroke-width:2px;

ThreadMetaGlobal is where we declare the static global state for all threads for the duration of the programme which can include multipel DB Instances.

ThreadMetaGlobalrust
// ---- Thread Static Meta ---- //

pub(crate) struct ThreadMetaGlobal {
    pub(super) thread_mu: Mutex<()>,

    pub(super) head: UnsafeCell<ThreadData>,

    pub(super) unref_handler_map: UnsafeCell<HashMap<usize, super::thread_local_ptr::UnrefHandler>>,

    pub(super) next_tls_id: AtomicUsize,

    pub(super) tls_id_free_list: UnsafeCell<Vec<usize>>,
}

Inside the meta global we have everything we need to manage the multiple threads in the matrix, we store the head of the linked list, with every link being ThreadData so it is an intrusive linked list. We also hold the map of unref_hanlders which tells the thread how to destruct a particular sub-system type. We also hold the next tls_id and a free_list of tls_id's that we can re-use.

We initialise ThreadMetaGlobal using std::sync::OnceLock basically the Singleton Pattern to ensure that we only create one per programme.

thread_meta()rust
pub(crate) fn thread_meta() -> &'static ThreadMetaGlobal {
    #[cfg(not(feature = "loom"))]
    {
        use std::sync::OnceLock;

        static STATIC_META: OnceLock<Box<ThreadMetaGlobal>> = OnceLock::new();
        STATIC_META.get_or_init(ThreadMetaGlobal::new)
    }
    #[cfg(feature = "loom")]
    {
        loom::lazy_static!(
            static ref STATIC_META: StaticMeta = StaticMeta::new();
        ) & STATIC_META
    }
}

This allows us to call thread_meta() anywhere in the codebase and be given the single ThreadMetaGlobal object and if there is none then the first call to thread_meta() will initialise it.

Of course, by being able to access it, we open ourselves up to concurrent risks if many threads are traversing or updating thread meta. For this we use a global mutex and atomics to avoid racing or deadlocks.

The second component is ThreadData.

ThreadDatarust
// ---- ThreadData ---- //

pub(super) struct ThreadData {
    pub(super) next: Cell<*mut ThreadData>,
    pub(super) prev: Cell<*mut ThreadData>,

    // Entries - columns in the thread local matrix, each column can comprise of multiple thread-local-storage sub-systems each with a unique tls_id
    entries: UnsafeCell<Vec<AtomicPtr<()>>>,
    registered: Cell<bool>,
}

ThreadData is the physical representation of one row in the matrix. The row belongs to a single thread, while each index in entries corresponds to a globally allocated tls_id and therefore to one logical subsystem column.

A new ThreadData is fully constructed but starts in an empty and unregistered state: entries has a length of zero, registered is false, and the intrusive-list links are null. The row is registered and populated lazily on first access through a ThreadLocalPtr.

The first-access path has three stages:

  1. ThreadData::ensure_registered() acquires thread_mu and links the row into the global intrusive list. Once linked, the row becomes discoverable by lifecycle operations that traverse all registered threads.
  2. ThreadData::with_tlp_ptr() checks whether the row contains an addressable slot for the requested tls_id. If entries.len() <= tls_id, it resizes the vector to tls_id + 1 and fills the new cells with null pointers.
  3. ThreadLocalPtr::get_or_init() inspects entries[tls_id]. If the cell is null, it invokes the subsystem's initializer and publishes the returned pointer into that cell. Later accesses by the same thread reuse the stored pointer.

Rows do not need to have equal physical lengths. The column identity is global, but each row materialises only the prefix required by the highest tls_id that thread has accessed. A thread that has never touched a particular subsystem may therefore have no slot for its column at all.

ThreadData registration, row growth, and entry initialization

The entries are type-erased because a single vector cannot directly contain the unrelated concrete types used by every subsystem. Each cell therefore stores an AtomicPtr<()>. The corresponding ThreadLocalPtr<T> acts as the typed capability for that column: it owns the tls_id, casts the erased pointer back to T, and registers the reclamation handler associated with that type and ownership protocol.

ThreadLocalPtr::new()rust
impl<T: ThreadLocalObject> ThreadLocalPtr<T> {
    pub(crate) fn new() -> Self {
        Self::new_with_handler(T::handler())
    }
}

ThreadLocalPtr<T> does not imply ownership of a Box<T>. A subsystem may reclaim its entry by freeing an allocation, decrementing a reference count, returning an object to a pool, deferring reclamation, or deliberately doing nothing. The ThreadLocalObject handler defines which protocol applies.

Safety Invariants

The matrix crosses Rust's normal ownership boundaries, so its soundness depends on a set of explicit invariants:

  1. Stable row addresses. Once a ThreadData row is linked into the intrusive list, its address must not change until it is unlinked. The thread-local allocation keeps the row at a stable address for the lifetime of its owning thread, while the boxed sentinel remains stable for the lifetime of ThreadMetaGlobal.
  2. Registry synchronisation. Every structural operation on the intrusive list, handler map, and TLS ID free list is performed while holding thread_mu. Registration, row teardown, and column reclamation therefore cannot mutate or traverse the registry concurrently.
  3. Single-threaded row access. Normal ThreadLocalPtr operations resolve to the calling thread's own row. Only that thread may initialise a cell or construct a Rust reference from its pointer. Cross-thread lifecycle operations handle entries only as raw pointers and never manufacture references from them.
  4. Stable column identity. While a ThreadLocalPtr<T> is live, its tls_id uniquely identifies that column. Every non-null pointer in the column must refer to an object compatible with T and with the handler registered for that ID.
  5. Valid entry lifetimes. NonNull<T> proves only that a pointer is non-null. The subsystem must ensure that the pointer is valid, correctly aligned, and remains alive until the cell is cleared or its handler reclaims it.
  6. No conflicting references. When get_or_init_mut() temporarily constructs &mut T, no other reference or access may conflict with it, and access to the same TLS entry must not be re-entered while the closure is running.
  7. Quiescence before column reclamation. Before a ThreadLocalPtr is dropped, its owner must guarantee that no thread can access that column again. thread_mu protects the registry traversal, but normal TLS access does not always acquire it and the mutex therefore cannot establish this quiescence by itself.
  8. Safe ID reuse. A tls_id is returned to the free list only after its old handler has been removed and every registered row has had that column cleared. A later subsystem reusing the ID must observe null cells rather than stale pointers from the previous owner.

The entries vector also has two valid access modes. The owning thread may access its own row while the subsystem lifecycle guarantees that no cross-thread reclamation is running. Any cross-thread structural access must instead occur under thread_mu. These conditions are what make the interior UnsafeCell<Vec<AtomicPtr<()>>> safe to use; the UnsafeCell itself provides no synchronisation.

Dropping A Thread Row

When a thread exits, Rust begins destroying its thread-local values and ThreadData::drop() calls drop_row(). At this point the owning thread must not perform any further TLS access.

drop_row() acquires thread_mu, unlinks the row from the intrusive list, and nulls its next and prev links. Unlinking first is important: once the row is removed, a concurrent column traversal cannot discover storage that is about to disappear.

The row then walks only the entries it has materialised. Each non-null pointer is atomically swapped to null before the handler associated with that index is invoked. Clearing the cell first ensures that another serialised reclamation path cannot reclaim the same entry twice. If no handler is registered for that column, the matrix clears the pointer but performs no object-specific reclamation.

Row teardown does not return any IDs to tls_id_free_list. A row represents one exiting thread, whereas each tls_id represents a column that may remain active in every other thread.

Dropping A TLS Column

Dropping a ThreadLocalPtr<T> is a different operation: it removes one subsystem column from every registered row. This is a global lifecycle transition and requires the owning subsystem to establish quiescence before ThreadLocalPtr::drop() begins. No worker may still call get(), init(), get_or_init(), or get_or_init_mut() for that column.

Once quiescence has been established, remove_column() acquires thread_mu, removes the column's handler from unref_handler_map, and walks the intrusive list. Rows whose vectors are shorter than tls_id + 1 are skipped because those threads never accessed the column. For every existing cell, the pointer is atomically swapped to null before the saved handler is invoked.

Finally, after all registered rows have been visited, the tls_id is returned to the free list. Holding thread_mu across this sequence serialises column removal with row registration and row teardown. Whichever teardown path reaches a cell first replaces it with null, so the other path observes an empty cell and does not reclaim the object again.