The runner was crashing every two and a half minutes, and one repo's job couldn't start for over a day. We were chasing a network bug—packet captures, WinDbg memory dumps, the works. The real cause turned out to be a single line in an API filter tangled up with the watchdog's own logic.
After decomposing the monorepo, each of the five repositories got its own CI. Architecturally beautiful, but brutally expensive in GitHub Actions minutes. We spun up our own Windows machine with a single runner for all five—and wrote a watchdog for it that would spectacularly fail less than 24 hours later.