Fixing a fork hang in tt-metal
A program that forked worker processes could not open a Tenstorrent device from a child process: the call never returned. We found out why and contributed a fix.
The problem
tt-metal is Tenstorrent's open-source framework for running models on their AI accelerators. Its maintainers had an open bounty issue for a hang: opening a device in a forked process, for instance through Python's multiprocessing, blocked forever.
Nothing crashed and nothing was logged. That is what makes this kind of bug expensive. It looks like a slow machine until somebody attaches a debugger.
The cause
fork() gives the child a copy of the parent's memory, but only the one thread that called it. In tt-metal, devices are managed by a process-wide singleton, the DevicePool, guarded by mutexes.
After a fork the child holds a copy of that pool and of its locks, frozen in whatever state they were in at that instant and describing devices the child never opened. The first device call in the child waits on a lock that no thread in the child will ever release.
The fix
POSIX has a hook for exactly this. pthread_atfork registers handlers that run just before a fork, and just after it in the parent and in the child. The patch registers three, the first time the pool is initialised:
- before the fork, take a lock so the pool is not mid-change when it is copied;
- afterwards in the parent, release it;
- afterwards in the child, drop the inherited singleton and re-create the lock, so the child builds its own pool the first time it opens a device.
void device_pool_child_postfork() {
// The child must not reuse the parent's pool. Clear it, so the
// first device opened in this process builds a fresh one.
if (DevicePool::instance_ptr() != nullptr) {
DevicePool::reset_instance();
}
new (&fork_safety_mutex) std::mutex();
}
pthread_atfork(device_pool_prefork,
device_pool_parent_postfork,
device_pool_child_postfork);It came to 60 lines across two files.
What happened to it
The pull request was closed without being merged. By then the maintainers had settled on a workaround that was serving them well, and making tt-metal properly usable from several processes needs deeper changes than they wanted to take on at that point.
They paid the bounty in full, and said they would build on the patch if they return to multi-process support. Both the issue and the pull request are public:
- The issue: cannot open device from a forked process
- The pull request: fix device hang when opening from forked process
What we take from it
- A hang with no log line is usually a lock. Start by asking who holds it.
fork()and libraries that hold locks or hardware handles do not mix by default. Code that may be forked has to either make its state fork-safe or say clearly that it is not.- In Python, starting workers with
spawninstead offorkavoids the whole class of problem, at the cost of a slower start.
Something of yours hanging, crashing or crawling? That is what project rescue is for.