Fixing a fork hang in tt-metal

A program that forked worker processes could not open a Tenstorrent device from a child process: the call never returned. We found out why and contributed a fix.

The problem

tt-metal is Tenstorrent's open-source framework for running models on their AI accelerators. Its maintainers had an open bounty issue for a hang: opening a device in a forked process, for instance through Python's multiprocessing, blocked forever.

Nothing crashed and nothing was logged. That is what makes this kind of bug expensive. It looks like a slow machine until somebody attaches a debugger.

The cause

fork() gives the child a copy of the parent's memory, but only the one thread that called it. In tt-metal, devices are managed by a process-wide singleton, the DevicePool, guarded by mutexes.

After a fork the child holds a copy of that pool and of its locks, frozen in whatever state they were in at that instant and describing devices the child never opened. The first device call in the child waits on a lock that no thread in the child will ever release.

The fix

POSIX has a hook for exactly this. pthread_atfork registers handlers that run just before a fork, and just after it in the parent and in the child. The patch registers three, the first time the pool is initialised:

  • before the fork, take a lock so the pool is not mid-change when it is copied;
  • afterwards in the parent, release it;
  • afterwards in the child, drop the inherited singleton and re-create the lock, so the child builds its own pool the first time it opens a device.
void device_pool_child_postfork() {
    // The child must not reuse the parent's pool. Clear it, so the
    // first device opened in this process builds a fresh one.
    if (DevicePool::instance_ptr() != nullptr) {
        DevicePool::reset_instance();
    }
    new (&fork_safety_mutex) std::mutex();
}

pthread_atfork(device_pool_prefork,
               device_pool_parent_postfork,
               device_pool_child_postfork);

It came to 60 lines across two files.

What happened to it

The pull request was closed without being merged. By then the maintainers had settled on a workaround that was serving them well, and making tt-metal properly usable from several processes needs deeper changes than they wanted to take on at that point.

They paid the bounty in full, and said they would build on the patch if they return to multi-process support. Both the issue and the pull request are public:

What we take from it

  • A hang with no log line is usually a lock. Start by asking who holds it.
  • fork() and libraries that hold locks or hardware handles do not mix by default. Code that may be forked has to either make its state fork-safe or say clearly that it is not.
  • In Python, starting workers with spawn instead of fork avoids the whole class of problem, at the cost of a slower start.

Something of yours hanging, crashing or crawling? That is what project rescue is for.

Let's Work Together

Tell us about your project — software, hardware, or both. New build, stuck project, or just exploring options — we're here to help.

Need It Fast?

Select your priority level and we'll respond accordingly.

Response Time
24–48 hours — faster for urgent projects
Free Consultation
No obligation, no pressure
Worldwide Service
Remote from Alberta, Canada

Get In Touch

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.