AI Foundry lab · Problems report · 25 September 2026, updated 27, 28, 29 and 30 September and 2 and 4 October

Requests for Roman: what trips people up on the AI Foundry ET-SoC-1 machines

Updated 4 October, 12:00 PDT: three updates of 1 October restored, two requests on die temperature, and the state after the visit. A rebuild of this page on 2 October overwrote the three updates of 1 October; they are back: the driver's counter race (C29, CF13; filed publicly on 4 Oct as et-platform issue #136), aifoundry1's disk, fixed on 30 Sep (H3, MO1, PO4), and the chip diagram's six asks (section 2.7, CD1–CD6; Policies and our own root steps are now sections 2.8 and 2.9). New: the cards report only a mean die temperature and a peak held since the card started, which our dashboard called the hottest sensor (C30), and CF14 asks for every sensor's reading and says how to get what exists today. Also new: C31–C33, H34–H37 and D25, the requests CF15, MO8, AS5 and SH9, and our root step U28. Changed since 2 October (read-only checks on all three hosts, 11:10–11:55; no card opened): the visit replaced card 0's fan but did not reach aifoundry2, whose card has been off the bus since 12:02 on 2 Oct with the host left on (SH5, H28); card 0's link logs about 7 corrected errors an hour again (C15, now partly fixed); the banners installed on aifoundry1 and aifoundry2 contradict the cards' state until a root install (H37); our once-a-second live monitor opens every card without its lock and fills two journals with sudo lines (H35, H9, H32); sudo grew on 2 Oct (AS2); and CF8 asked for a reading no command gives, so it is now a short test. The visit list below is now the next visit's, and host addresses and a root login recipe are gone from H2, H23 and C28.

For the next lab visit (not yet scheduled): tasks in order (open “Detailed instructions” under each). The visit of Friday 2 October replaced aifoundry1 card 0's fan (task 2); the other tasks are still to do.

  • 1. aifoundry2's card cooling: the BIOS fan settings (SH5)

    Urgent, and not done on 2 October: the card fails within about 70 minutes of every boot. On 2 October we watched it twice. After the full reset at 10:53 the idle card went from 45 °C to 85 °C in 22 minutes, crawled to 94 °C over the next half hour, then ran away: 105 °C at 11:58, 126 °C at 12:01:31 and 138 °C (peak sensor reading 144 °C) at 134 W at 12:02:01, when it dropped off the bus (the same as 07:42 that morning, at 126 °C and 103 W). The machine was idle and nobody used the card. Our fit of the 828 readings: the card sheds its heat into air at about 51 °C, against about 31 °C for aifoundry3's identical card, and it would settle at about 81 °C with air just 3 °C cooler. So the fault is the air reaching this card (its own fan, a fan header that stops, or no airflow in that corner of the case), not the chip. The card must not run until the cooling is fixed. It has been off the bus since 12:02 on 2 October, and aifoundry2 has stayed on since then with the card in it; the host's own drive, network-chip and CPU sensors have tracked aifoundry3's within about 2–3 °C, so there is no sign that it still heats. Powering aifoundry2 off until the visit is the owner's call (it also runs our dashboard and history jobs). First the card and its fan (the first steps), then the fan settings, then a 30-minute test with us watching. About 60 minutes; needs a monitor and keyboard.

    Detailed instructions
    1. First, the card itself: shut aifoundry2 down (sudo poweroff), switch the power supply off at its rear switch and unplug it. Look at the ET card (the second long slot from the CPU; the top slot is empty): does it have its own fan, and does it turn freely by hand? Is the heatsink seated flat and fully screwed down, with no dust packed in the fins? Compare it with aifoundry3's card, which runs cool. If its fan is dead or stiff, the card needs a new fan or heatsink before it runs again (AI Foundry's card: tell Roman).
    2. Follow the card's fan cable, if it has one. If it plugs into a motherboard fan header (SYS_FAN…) rather than the card itself, that header must run at full speed: in the BIOS step below, set that header to Full Speed, never a curve that follows the CPU. A fan that stops when the CPU idles is exactly what our data show.
    3. If the card looks fine, reseat it: unplug its PCIe power cable(s), unscrew its bracket, press the slot latch, pull it straight out, push it back in until the latch clicks, screw the bracket, and plug the power cable(s) back in fully.
    4. If you will also update the BIOS (task 3), do that on the machine next: a BIOS update resets every setting below.
    5. Tell the lab first: shut each machine down cleanly with sudo poweroff (or a short press of the power button), not by holding the button.
    6. aifoundry3: power on and press Del repeatedly to enter the BIOS. Press F2 for Advanced Mode if it opens in Easy Mode. Press F6 (or click Smart Fan 5). For each fan header in the list (SYS_FAN1 … SYS_FAN6, CPU_FAN, CPU_OPT), photograph the screen: Fan Speed Control, Temperature Input, Fan Stop, and the curve. Also photograph PC Health Status (the fan speeds). Press Esc, choose Exit without saving.
    7. aifoundry2: enter the BIOS the same way, then F6 for Smart Fan 5. For every SYS_FAN header: set Fan Stop → Disabled; Fan Speed Control → Manual, and drag every point of the curve to at least 50% (or choose Full Speed if noise does not matter); if Temperature Input offers anything other than CPU (System 1, System 2, PCH or a PCIe slot sensor), choose that; Fan Fail Warning → Enabled. If aifoundry3's photos show different values that keep its card cool, copy those instead.
    8. Still in the BIOS: open PC Health Status and write down every fan's RPM. A fan that shows 0 RPM or N/A while the CPU fan spins is stopped or dead: note which header.
    9. Settings tab → Platform Power → AC BACK → Always On (the machine starts by itself after a power cut), and ErP → Disabled.
    10. Press F10, check the list of changes it shows, choose Yes to save and exit. The machine boots Ubuntu.
    11. With the case open and the machine running, check which fans blow across the ET card and that they all spin. If no fan reaches the card, fit a case fan or a slot fan bracket aimed at it.
    12. The test, with us watching (the card's temperature now streams once a second: the dashboard's Live section, or the History page, Hour view). Boot with the side panel off and tell us. It must level off below about 70 °C within 20 minutes with the host idle (aifoundry3's identical card idles at 52–60 °C; the same pass mark as SH5). If it passes 95 °C it will run away within about 10 minutes: switch the machine off.
    13. If it still climbs, a quick airflow test: point any desk fan into the open case at the card and watch for 15 minutes. If it levels off, the fix is airflow: fit a case fan or a slot fan bracket aimed at the card.
    14. If even that fails, swap the cards of aifoundry2 and aifoundry3 (same board, same slot; power both off and unplug them first). If the overheating follows the card to aifoundry3, the card's own cooler is faulty and the card goes back to AI Foundry; if it stays with aifoundry2, it is that machine's airflow.
  • 2. Done: aifoundry1 card 0's fan was replaced (SH1)

    Done on 2 October: its fan was found broken and replaced. In our test that afternoon it ran cooler than card 1 (52–53 °C, peak 56 °C, under 8 minutes of sgemm), and it is back in service; on 4 Oct it idled at a 48–55 °C mean. Nothing is left on site. From here: aifoundry1's login banner still warns against card 0 until the corrected one is installed as root (U28), and its link's corrected errors, about 7 an hour since that boot, are watched (C15). The steps below are kept for the record.

    Detailed instructions
    1. Fan settings: the same steps as task 1 on aifoundry1 (photograph its current settings first).
    2. Shut down (sudo poweroff), switch the power supply off at its rear switch, unplug it, and touch the bare metal of the case to discharge static.
    3. Compare the two ET cards: does card 0 have a fan on its heatsink, and does it spin freely by hand? Is its heatsink flat and fully screwed down? Is dust clogging its fins? Card 1 is the reference.
    4. Reseat card 0: unplug its PCIe power cable(s), unscrew its bracket, press the slot's release latch, pull the card straight out, check the gold contacts for dust, push it straight back in until the latch clicks, screw the bracket, and plug the power cable(s) back in fully (they click).
    5. Do not change the thermal paste or open the heatsink without Roman's OK: it is AI Foundry's card.
    6. Power on and tell us: we run ten minutes of load on card 0 remotely and compare it with card 1. If it still overheats, the card goes back to AI Foundry or gets a new heatsink.
  • 3. Optional: the BIOS update on all three (SH3)

    All three run 2021 BIOSes (F5 on aifoundry1 and aifoundry2, F6 on aifoundry3); the current release is F9 (June 2023). About 15 minutes per machine; it resets every BIOS setting, so do it before tasks 1 and 2 on that machine. Needs one USB stick.

    Detailed instructions
    1. Download mb_bios_z590-aorus-master_f9.zip (10 MB; also staged on aifoundry2 in ~/claude/private/visit/). Format a USB stick as FAT32 and unzip the archive onto the root of the stick, so that flash.nsh, Efiflash.efi, Z590AORUSMASTER.F9 and the EFI folder sit at the top.
    2. Shut the machine down cleanly, plug the stick in, power on and press F12 for the boot menu; choose the USB stick (the UEFI entry). It boots into an EFI shell.
    3. At the Shell> prompt, find the stick: it is usually FS0: (the list printed at start shows which FSn is the removable one). Type FS0: and Enter, then flash.nsh and Enter.
    4. The tool checks the BIOS and restarts the machine. Leave the stick in and do not switch the machine off. After the restart the update runs by itself (a few minutes), then the machine restarts again.
    5. Enter the BIOS (Del) and check the version on the main screen says F9. Then redo: task 1's fan settings and AC BACK; and check Boot → Boot Option #1 is ubuntu (on the Samsung SSD). Save with F10.
    6. If Ubuntu does not start, press F12 at power-on and choose ubuntu, then set it as Boot Option #1 in the BIOS.

    We re-check the hosts afterwards (driver, cards, timings): a new BIOS changes host-side timing, so measurements before and after are not compared.

  • 4. Optional: the SSD firmware and the driver key (SH3)

    All three boot from a Samsung 980 PRO on firmware 3B2QGXA7, which has a known health bug; the fix is 5B2QGXA7. We can apply it from Linux as root later, after a backup of aifoundry1 (it has none), so no visit is needed. The driver key matters only before anyone turns Secure Boot on.

    Detailed instructions
    1. SSD, on site if you prefer: download Samsung_SSD_980_PRO_5B2QGXA7.iso (28 MB; also on aifoundry2 in ~/claude/private/visit/), write it to a USB stick with balenaEtcher or sudo dd if=Samsung_SSD_980_PRO_5B2QGXA7.iso of=/dev/sdX bs=4M, boot it with F12, and answer Y when it offers 5B2QGXA7; then power off and on.
    2. The driver's signing key (only if Secure Boot will be turned on): on the machine, sudo mokutil --import /var/lib/shim-signed/mok/MOK.der, type a one-time password twice, and reboot. On the blue Perform MOK management screen choose Enroll MOK → Continue → Yes, type the same password, then Reboot.
  • 5. Hardware: Ethernet, memory, a remote console, power (SH4, SH7, SH6, SH2)

    All three are on Wi-Fi only; aifoundry3 runs on one memory stick; nobody can reach a console remotely; and power was cut to all three at once in July and on 18 Sep (and on purpose, by Roman, on 30 Sep after C28).

    Detailed instructions
    1. Ethernet: plug a cable from the lab's switch or router into each machine's Ethernet port on the back panel. Nothing else is needed: Ubuntu's existing “Wired connection 1” comes up by itself and Wi-Fi stays as the fallback. We confirm it from here.
    2. Memory for aifoundry3: buy one 32 GB DDR4 desktop DIMM (DDR4-2666 or faster; it will run at 2666 to match the one there). With the machine off and unplugged, install the pair in the slots the board's label or manual names for two DIMMs (on Gigabyte Z590 boards: DDR4_A2 and DDR4_B2, the second and fourth slots from the CPU), moving the existing stick if needed. The first boot after a memory change can take a minute of black screen: wait.
    3. Remote console: ask Roman for an IP-KVM (such as a PiKVM) on at least aifoundry1, or the name and phone number of someone who can press a power button during working hours.
    4. Power: ask what cut power to all three machines at once (17 and 23 July, 18 Sep): a breaker, a shared power strip, a building outage, or someone switching them off. Ask for a UPS for the three (about 1500 VA covers three idle machines) and that nobody power-cycles the shared strip to reset one machine.
  • 6. Decisions to ask Roman for

    Each is a question for Roman; the request linked gives the background.

    Detailed instructions
    1. Thermal protection and one firmware release (CF1, PO1): “Can the firmware team say which lab release has the three upstream fixes, and can the cards get a die-temperature cut-off? Can all four cards run one release?”
    2. Root and sudo (AS1, AS2): “Any account on one lab machine can log in as root on the others, 18 accounts on aifoundry2 have sudo (4 Oct), and since 2 Oct newcomers create their accounts from the shared root login. Who should keep root and sudo, and how should accounts be made without root (AS5)?”
    3. aifoundry1's disk (MO1, MO2): “aifoundry1's disk is no longer full (116 GB free since 30 Sep), but it has four corrupt files in other users' data and a toolchain tree, and no backup. Can the owners restore or delete those files, and can the lab back it up?”
    4. The card lock (PO2): “Can using the card lock (flock /run/lock/etsoc-shire<N>.lock) be the lab's rule for everyone, including the CI runners and aifoundry3's demo?”
    5. The New user page (DI1): “Can the lab adopt our public New user? Start here page and link it from #community-lab? The login banners link it since 30 Sep, and its first step follows the lab's 2 Oct rule.”
    6. aifoundry3's admin (DI4): “Please tell aifoundry3's admin about our changes there on 30 Sep (headless boot, desktop services off, peer names, et-reset and a daily health check), and since then the card-usage logger, et-lab-start, our once-a-second live monitor, and accounts newcomers create themselves.”
Where things stand, and what this page is

In two and a half weeks of measurements on the lab's four ET-SoC-1 cards we hit 87 problems that will trip up other users too, and 8 more in our own tools and docs (C26, H30, D22, D23, D24, H34, H35, H37): two of the four cards could not be opened by any ET tool for at least a week (and once they could, one of them overheated under load), the cards differ in firmware, clock policy and thermal protection in ways the driver does not show, and the hosts reboot together without warning, run out of disk, and until 25 September hid their kernel logs from users. On 25 September, with root at the owner's request, we fixed 12 of them on the machines (and D16 in our own deploy script), and all of those still hold. On 30 September our own link test hung aifoundry1 until a power cycle (C28). Of all 95, 25 are fixed, 8 partly fixed, 40 open and 22 new and not yet fixed. 58 of the 70 not yet fixed need the Nekko team: a decision, a policy, the site, the firmware, the driver or the runtime; the other 12 are ours to close (sections 2.9 and 4.4).

State on 4 October, 12:00 PDT. Three of the four cards work. aifoundry1's card 0 is back in service since its fan was replaced on 2 Oct, idling at a 48–55 °C mean (C21); its root port counts about 7 corrected errors an hour (C15). aifoundry2's card has been off the PCIe bus since its idle runaway at 12:02 on 2 Oct, and the host has stayed on with the card in it; its cooling is the next visit's first task (SH5, H28). The login banners installed on aifoundry1 and aifoundry2 still describe the cards as they were before 2 Oct, until a root install (H37, U28). The driver's counter race was filed publicly upstream on 4 Oct, as et-platform issue #136 (C29). Of all 95 problems, 25 are fixed, 8 partly fixed, 40 open and 22 new.

What needs the Nekko team is in section 2, Requests for Roman: 54 requests in 8 groups, each sorted by importance; what needs only root on the machines, and no one else's decision, is ours (section 2.9). Beyond the visit list above, the five that matter most: real thermal protection and one firmware release (CF1, PO1), root between the machines, who keeps sudo, and accounts without the shared root (AS1, AS2, AS5), aifoundry1's corrupt files and a backup (MO2), the power cuts (SH2), and the card lock for everyone (PO2).

Earlier updates: 2 October, 30, 29, 28 and 27 September

State on 2 October, 14:45 PDT. Three of the four cards work. aifoundry1's card 0 is back in service: its broken fan was replaced on site, and it now runs cooler than card 1 (C21, SH1). aifoundry2's card overheats whenever it is on: after each full-reset boot today (06:45 and 10:53) it heated with the host idle until it dropped off the PCIe bus, at 07:42 and at 12:02 (138 °C, 134 W: a thermal runaway, watched once a second, H28). Its cooling is the visit's first task (SH5). On 30 September our link test hung aifoundry1 until Roman power-cycled the lab (C28); since that cold boot its card 0 has logged no link errors (C15). The root steps that wait on no one else are done on all three hosts (U16, U18, U20, U22, U24, U27). Of all 86 problems, 25 are fixed, 7 partly fixed, 42 open and 12 new.

State on 30 September, 15:45 PDT. At 14:41 our test of aifoundry1 card 0's PCIe link at 8 GT/s, run as root by the owner, hung the whole host (C28); Roman power-cycled the lab at about 15:07, and all three hosts came back on kernel 7.0.0-34 with every card working, which also completed the pending reboots (H6) and proved the headless boot (H8) and the driver's load at boot (C2). The cause is not known yet; the lesson is on its own page. Earlier that afternoon the owner ran the root steps of section 2.9 that wait on no one else: aifoundry2 now boots headless, has the desktop services and apport's hook off, takes no password logins and finds its peers by their tailnet addresses (U27, U20, U22, U24); aifoundry3 has the same, apart from the logins (it has no sshd) and one wrong line in its peer names, and it and aifoundry1 now have et-reset and a daily health check (U16, U18). A read-only re-check of every problem that afternoon found nothing else changed on the machines, and D22 fixed since 27 September. Of all 85 problems, 22 are fixed, 8 partly fixed, 43 open and 12 new.

Updated 29 September, 01:10 PDT: four owner-approved fixes on aifoundry1, and our own clean-up. On 28 September at 20:51 PDT, with the owner's approval, four items of section 2.9 ran on aifoundry1: the headless boot target, from its next boot (U27); bluetooth, cups-browsed, the firmware-updater snap and apport's coredump hook switched off (U20); one login-service setting (U22); and the other two machines pinned to their tailnet addresses in /etc/hosts (U24). Each item now says what changed and how to roll it back, and the host keeps the session log and the files as they were before (section 4.3). On aifoundry1 they fix the password half of H19, H23 and most of H33, and H8 from its next boot, so H19, H23 and H33 are now partly fixed. The same items on aifoundry2 and aifoundry3, and the other eight items, did not run and wait for the owner: this session's permission check refused the ones we tried that evening, among them the scrub (U17) and the Gen3 test (U25). As our own user we also removed a crash report of ours from aifoundry1's /var/crash (C19), masked the firmware notifier and wireplumber in our user managers on all three hosts (U11, now done), rebuilt aifoundry1's build/sparsity with the g3log fix (C17), and replaced the stale run_pcie.sh on aifoundry1 and aifoundry3; the fixed script's first card run, at 23:56, released the card lock between sub-tests (D24, now fixed). A read-only check at 01:10 on 29 September found every change in place. Of all 84 problems, 17 are fixed, 11 partly fixed, 43 open and 13 new.

Updated 28 September, 08:52 PDT: aifoundry2's card runs kernels again, and what needs only root is ours now. The per-card sysfs reset at 06:39 had re-attached the hung card but left its Master Minion hung (launches at 06:41–06:47 still failed); the management reset at 08:32 recovered it, and a test kernel ran three launches at 08:33 (C27, now fixed; C14). With the owner's approval we also destroyed our 25 September snapshots on aifoundry1 at 08:33, which freed about 2.4 GB (U3, H3), and stopped our three orphaned Chrome processes on aifoundry2 at 08:34 (U11). The owner has root on all three machines, so everything that needs only root there, and no one else's decision, has moved out of the requests for Roman into our own list (section 2.9): 2 items done and 12 waiting for the owner's decision, each with its exact command and its risk to other users. 41 requests remain for the Nekko team and the lab; a table says where each of the 47 went.

Updated 28 September, 08:02 PDT: the root steps are done, at the owner's request. With root again from 06:30, we exported aifoundry3's journal back to 17 July and raised the 2 GB cap we had set on 25 September to 4 GB (H32, now fixed; between the export and the cap change journald deleted its two oldest files, of 17 July, which survive only in the export); installed et-who --check, the corrected banners and et-lab-health on all three hosts (H30, now fixed); read every card's PCIe payload sizes and the hosts' memory as root: MaxPayload 256 B and MaxReadReq 128 B on every card (D19, H29); and on aifoundry1 cleared systemd's reload warning and compressed the flood-era logs (C16). Each step was dry-run first and backed up under /root/labfix-20260928/, and a separate read-only check at 08:00–08:02 confirmed every change; none of them opened, queried or reset a card. Two steps did not run then, because this session's permission check refused them: destroying our ZFS snapshots on aifoundry1 (U3, H3) and stopping our orphaned Chrome on aifoundry2 (U11); both ran later that morning with the owner's approval (above). Details in section 4.3.

Updated 28 September, 03:30 PDT: aifoundry2's card hung. Its Master Minion hung at 02:50:53 PDT during our DVFS development run (C27): the service processor still answered, but no kernel could run on the card until it was reset, which held up the owner's DVFS validation. Following the lab's rule, we did not try a reset overnight. It was recovered at 08:32 with the owner's approval: the per-card sysfs reset at 06:39 had re-attached the card but not its Master Minion, and the management reset worked (above). The cause is not established; CF3 asks the firmware team to look at our hypothesis, a kernel launch during the governor's idle clock reset. Also new: in one of our 26 September campaign passes aifoundry2 ran at a 90–103 °C mean for 4 minutes with nothing tripping, more evidence that no card has a die trip (C22); that pass also broke our own 90 °C rule.

Updated 27 September, 22:30 PDT; corrected at 23:40. We re-checked all three machines (read-only, 21:26–21:45 PDT) and went through everything measured since 25 September. Of the 66 problems of 25 September, 13 are fixed and all of them still hold, 8 are partly fixed (among them C2, counted as fixed on 25 September but untested on aifoundry1 until its reboot), and 45 are still open; 12 of the open ones have new facts, and one of those, aifoundry1's disk, was worse (H3; about 2.4 GB of it, our snapshots, came back on 28 September). 18 problems are new (C22–C27, H28–H33, D19–D24; C27 was added on 28 September; C27, H30 and H32 were fixed that morning). The most serious: no card has a die-temperature hard trip (C22), aifoundry3's clock pin also stops its governor (C23), aifoundry1 card 1's clock never moves (C24), and other accounts now work on aifoundry3, one of them on its card (H31).

Done on 27 September, as our ordinary user: we copied our working files out of aifoundry2's /tmp (H22; work continues there, so the copy is repeated before any reboot); moved a crash report of ours out of aifoundry1's /var/crash, where C19 had regressed; saved aifoundry3's wtmp boot history (H32); found that aifoundry3 runs on a single memory channel (H29); staged a new et-who, a health check and corrected banners on every host; corrected our repository docs and pushed them; and drafted the onboarding page and the per-card sheet (appendices A and B). The root steps were refused that night; most of them ran on 28 September (sections 4.2 and 4.3).

This page is public. It describes the open security findings (H2, H19) only as far as the lab needs to act on them.

1. All 95 problems at a glance#

Status on 4 October: notes dated 4 October say what changed since 2 October (read-only checks on all three hosts, 11:10–11:55 PDT; no card was opened), and problems added that day are marked New. Status on 30 September: every problem was re-checked read-only that afternoon (14:17–14:33 PDT), before the owner's root session (14:37–14:41) and aifoundry1's fall (C28); notes dated 30 September say what changed, and the others were unchanged. Earlier states: 27 September, 22:25 PDT, and the changes of 28 and 29 September: Fixed done, verified, and still holding on 27–30 Sep (with the date it was fixed); Partly fixed part of it is fixed, on some hosts or one side of it; Open nobody has fixed it; New found since the 25 September report and not yet fixed (C27, H30, H32 and D24, also new, were fixed on 28 Sep, and D22 on 27 Sep, and show as Fixed; H33, also new, is partly fixed and shows as Partly fixed); Regressed a fix of 25 September undone (none now: C19 regressed on aifoundry1 on 26 Sep and was cleared again on 27 Sep). Tags: updated what we know changed materially since 25 Sep, worse it got worse, workaround a known workaround avoids it. The line under each title is the latest evidence (27 September–4 October). Next names the requests that would close it (section 2) and our own steps still to do (sections 2.9 and 4.4). Fixed problems are greyed out but kept. Several problems were reported three times from different angles (cards, hosts, measurement); each appears here once.

Next
Cards, driver and firmware
C1aifoundry1's two cards were refused by every ET tool: the driver module had an empty version stringHolds through the 30 Sep and 2 Oct boots: all three hosts load the driver with its version string (0.20.0, the same srcversion), and aifoundry1's card nodes came up at the 2 Oct boot. What is left is the lab's: package the driver so it cannot recur, and a clearer error upstream.aifoundry1blocks workFixed 25 SepNekko team CF4, CF5; us U18
C2Nothing reliable loaded the ET driver at boot on aifoundry1 and aifoundry3Tested at boot on all three: after the power cycle of 30 Sep at about 15:07 each host loaded et_soc1 0.20.0 by itself and created every card's nodes, aifoundry1 for the first time since its fix of 25 Sep.aifoundry1, aifoundry3corrupts resultsFixed 30 SepNekko team CF4, DI4; us U19
C3Stale and foreign ET driver builds sat in DKMS, a risk at every kernel updateHolds. One install recipe (the package) is still the lab's.all threewastes time latentFixed 25 SepNekko team CF4; us U18
C4Four cards run three firmware releases and idle at different operating points; the driver reports the same nameplate for allNo sign of a reflash: /opt/et and the driver are unchanged (the firmware query was not run, since it opens the card). The dates of the /dev/et* nodes, which earlier checks used, are no evidence either way: they change at snap refreshes (aifoundry3 at 07:17 and aifoundry2 at 22:43 on 2 Oct, with nothing but snap and AppArmor in the kernel log). The kernel log is the check.all four cardscorrupts resultsOpen updatedNekko team PO1, DI2, DI3
C5aifoundry3 is held at 600 MHz by a root boot service, not by its firmware, and the host cannot see itSame pin, new consequence: aifoundry3 has no working thermal governor (C23).aifoundry3corrupts resultsOpen updated workaroundNekko team CF2, DI2
C6After a card reset, aifoundry3's clock guard trusts its boot marker and does not re-check the cardUnchanged: the guard script is as it was on 23 Jul, and aifoundry3's card has not been reset since 25 Sep 16:22. et-reset is installed on aifoundry3 since 30 Sep 14:39 (U16); it never touches the guard and prints a reminder to tell aifoundry3's admin instead, so after a reset the 600 MHz pin still has to be restored by that admin (CF2, DI4).aifoundry3corrupts results latentOpenNekko team CF2, CF3, DI4; us U16
C7On cards with a free governor the clock follows the die temperatureaifoundry2's governor follows the 34-sensor mean with no dead band (DV2's validation, 28–29 Sep), but its die idles above 65 °C, so it rarely acts (H28); card 1 never steps (C24), aifoundry3 is pinned (C23), and card 0 was not measured.measured on aifoundry2 only (from a cool die); aifoundry1 card 0 unmeasured, but it throttled 10 times on 25 Sepcorrupts resultsOpen updated workaroundNekko team CF3, CF8, SH5
C8The firmware on the cards is a mid-2024 build, but the source everyone reads is from December 2025We mapped all four cards to their source, and the mapping, with each build's governor, has been in the repository since 28 Sep (docs/findings/14-card-behaviour.md); 1.4.1's exact source is still not public (CF9).all four cardscorrupts resultsOpen updated workaroundNekko team CF9, PO1
C9Users cannot run changed firmware: the images must be signedUnchanged. The firmware changes the measurements need are rungs 11–16 of the hub's ladder.all four cardsblocks workOpenNekko team CF12
C10A management tool killed mid-request poisons the card's management queue, and samplers often fail to startNo incident since 25 Sep 17:00, and no vendor fix; the banners keep the never-kill-9 rule and the drain line.all threeblocks workOpen workaroundNekko team CF5, DI1
C11Nobody could see who held a card node, which only one process can openHolds. Its follow-up H30 was fixed on 28 Sep (et-who --check, installed on all three); C26 remains. Upstream: name the holder, and a read-only telemetry path.all threewastes timeFixed 25 SepNekko team CF5, CF6, PO2; us U13
C12Any user can change a shared card's global state: reset it or change its trace levelWider than reported: stats reset, thresholds, DVFS-off, TDP and frequency are open to every user too, and, from the firmware source (read for E54 on 28 Sep, untested), so is a rail's voltage: on card 1's 1.2.0 each NoC voltage set rewrites the flash sector the boot voltage comes from, and on the two 1.3.1 cards a failed set loops until the watchdog resets the card (CF10, PO1). Since 28 Sep the banners on all three hosts list the other commands (U5), not yet the voltage set.all threecorrupts resultsOpen updatedNekko team CF10
C13On the two-card host the stock ET tools open both cards; picking one needs aifoundry1's forked runtimeNew cost, from our own tools: a stock tool for card 1 collides with anything on card 0.aifoundry1blocks workOpen workaroundNekko team CF7, DI4; us U13
C14A hung card is recovered by power-cycling the whole host, although a per-card reset exists and worked on an idle cardNo hung card since 28 Sep (C27): the restored card ran DV2's 52 validation launches (28–29 Sep) without one. A checked reset command, et-reset, is installed on all three hosts since 30 Sep (U16; aifoundry2's at 15:36), for root and the sudo group. aifoundry2's card is now off the bus for another reason, its cooling (H28, C32): no reset can help it, and none should be tried before the cooling is fixed. Which reset is the supported recovery is still CF3's question, and who may run it the lab's (PO2).all threeblocks workOpen updatedNekko team CF3, PO3; us U16
C15aifoundry1 card 0's PCIe link logs about one corrected error per secondThe flood has not come back, but the count is not zero. Since aifoundry1's boot at 12:59 on 2 Oct (after the fan visit) card 0's root port has counted 336 corrected receiver errors in 46.7 hours, about 7 an hour, against about 3,600 an hour before 30 Sep; card 0 itself and card 1's port count none, and both links run at 16 GT/s x8. The kernel logs them at 2–17 lines an hour. Our dashboard shows this port's rate as 0 an hour, which is wrong (ours to fix). Watch it; on site only if the rate grows.aifoundry1 card 0wastes timePartly fixed updatedus watch it (U18), fix the dashboard's rate; on site only if it grows
C16The PCIe error flood filled aifoundry1's logs on a nearly full diskHolds. On 28 Sep we cleared the systemd reload warning (daemon-reload: 0 units need one) and compressed the flood-era kern.log.1 and syslog.1 (287 and 299 MB to 11.7 and 13.6 MB, content checked). The size cap stays off on purpose.aifoundry1wastes timeFixed 25 SepUs U18
C17Host programs on aifoundry3 crash about once in 100 launches, 1.08 s after startFixed in our programs on aifoundry2 and aifoundry3 (26 Sep: 641 processes, no crash); aifoundry1's were rebuilt on 28 Sep (six at 07:52, build/sparsity at 20:56), all but the campaign's enercat_v2; the runtime is unchanged, and another account now runs programs on aifoundry3 (H31).aifoundry3wastes timePartly fixed workaroundNekko team RT1, DI1; us U13
C18Each host runs a different build of the vendor runtime in /opt/etUnchanged: the three runtime builds still differ; E50 now shows that host-side timings follow each build.all threecorrupts resultsOpenNekko team RT2
C19The vendor tools mislead when something is wrongTwo more misleading readouts found (the maximum temperature reads 0; “low_power” is only a power threshold). apport's coredump hook, which wrote duplicate crash reports of our programs, is off on all three hosts since 30 Sep (aifoundry1 since 28 Sep; U20). On aifoundry2 the hook's failed unit cleared with the reboot, and the host reads running (4 Oct); only apport's own crash report of 28 Sep is left in /var/crash, for root to remove (U20).all threewastes timeOpen updatedNekko team CF5, CF1; us U20
C20The driver's signing key is not enrolled: turning Secure Boot on would make the cards disappearNew detail: the firmware is in Setup Mode, so a BIOS reset that restores the default keys could turn Secure Boot on and hide the cards.all threeblocks work latentOpenNekko team SH3
C21aifoundry1's card 0 overheats under load: 98–102 °C in short test runs, 115–117 °C just afterFixed on 2 October: on site, card 0's fan was found broken and replaced (host up at 12:59). Our acceptance test that afternoon: 49 °C idle (the peak since the restart 53 °C) against 65 °C before, and 8 minutes of sgemm bursts (8 s under the lock, 2.5 s gaps) held it at 52–53 °C (peak 56 °C), cooler than card 1 under the same test (59–60 °C, peak 63 °C). The owner put it back in service, and the new-user brief now offers it (SH1). 4 Oct: it has idled at a 48–55 °C mean for two days (peak 57 °C), at 19–21 W, with no error events. The “123 °C hot spot” quoted before was a peak held since September (C30). The installed login banner still warns against card 0 until the corrected one is installed as root (U28).aifoundry1 card 0blocks workFixed 2 Oct updatedus install the corrected banner (U28, root)
C22No card has a die-temperature hard trip or any hot-spot protection, and the lab's firmware releases carry known thermal and power bugsFound on 27 Sep in the firmware source of all three releases, and seen in our 26 Sep data (a pass at a 90–103 °C mean, up to 86.9 W, with the clock never leaving 600 MHz). 2 October: an idle card heated to 138 °C (peak sensor reading 144 °C) at 134 W, and nothing on the card acted; it dropped off the PCIe bus at 12:02 (H28). At idle a clock cut cannot help: the power is leakage, which grows with temperature (about doubling every 23 °C above a 13.6 W floor). Only a hard trip that lowers the minion voltage or cuts the minion rails, at about 105 °C, would stop a runaway; until the firmware has one (CF1), a host whose card loses its cooling has no protection.all four cardsblocks workNew workaroundNekko team CF1, PO1; us U1, U13
C23aifoundry3's 0 W TDP pin also stops its governor, so the card makes no thermal stepInferred from the source; every service-processor trace on aifoundry3 since 25 Sep is empty, as predicted.aifoundry3corrupts resultsNew workaroundNekko team CF2, DI2
C24aifoundry1 card 1's clock never moves: its DVFS appears to be off, so it has no thermal step either600 MHz in all 359,657 campaign samples, cool or hot, and again in all 10,033 samples of our 29 Sep runs; its trace probe was silent (27 Sep 20:24). 4 Oct: no command reads the active-power-management flag (the management API has only the set, which changes the card for everyone), and 1.2.0's closest public source (da192816a) turns it on at every service-processor boot with no VMIN check, so neither hypothesis below explains a card that never steps. Card 1's service processor restarted at the 2 Oct boot; a busy run on it that afternoon (13:07–13:15, 54–60 °C) logged no clock. One short cool-start run that logs the clock settles whether it steps now (CF8).aifoundry1 card 1corrupts resultsNew workaroundNekko team CF8, PO1, DI2
C25The operating system cannot see a card's temperature, or any chassis fan: only the single-opener management node reports itChecked on 27 Sep and again on 4 Oct on all three hosts: no ET entry in hwmon, no fan readings. Since 2 Oct the lab dashboard's Live section and the History page show each card's mean temperature once a second, read by our live monitor through the management node (H35).all threewastes timeNew workaroundNekko team CF6
C26Our own tools opened aifoundry1 card 0's management node without card 0's lockThe heat guards have not run since 28 Sep. But since 2 Oct our live monitor opens every card's management node, card 0's included, about once a second without the card's lock, while et-who --check shows no holder (H35).aifoundry1 card 0wastes timeNewUs H35, U13; Nekko team CF6
C27aifoundry2's Master Minion hung on 28 Sep; the sysfs reset did not recover it, the management reset didRecovered on 28 Sep: the per-card sysfs reset at 06:39 re-attached the card but left the Master Minion hung (launches at 06:41–06:47 still failed); the management reset at 08:32:45 recovered it, and a test kernel ran 3 launches at 08:33. It has not hung since: DV2's validation ran 52 launches on it on 28–29 Sep, all returning 0, one of them meeting the same clock step down as the hung launch. The cause, and which reset is the supported recovery, are CF3's questions.aifoundry2blocks workFixed 28 SepNekko team CF3
C28Retraining aifoundry1 card 0's PCIe link to 8 GT/s took the whole host down on 30 Sep; it needs a power cycle on siteOur Gen3 test of card 0's link (U25) retrained its root port to 8 GT/s at 14:41 on 30 Sep, and the whole host froze at once: the journal's last entry is at 14:41:17, with no kernel error, machine check or panic record. Roman power-cycled the lab at about 15:07, and it came back with both cards working (the incident and its lesson).aifoundry1blocks workFixed 30 Sepus U25 (never repeat)
C29A race in the ET driver: reading a card's message counters while the card is reset, or while the driver loads, can return garbage or crash the readerFound on 30 Sep by reading the driver's source while reviewing our new usage logger; never seen on a card, and not tried. The driver shows these counters before it builds the queue tables they read, and frees the tables before it removes the counters. Our usage logger already ignores an impossible jump. 4 October: filed upstream, publicly, as et-platform issue #136. The chips are end-of-life, so the owner chose a public report. et-platform's head is still 836a4ab of 17 Jul, and the driver on all three hosts is unchanged (0.20.0, the same srcversion). A patch is ready for the maintainers (CF13).all threecorrupts results latentNew workaroundNekko team CF13
C30The cards report one die temperature: the 35 sensors' own readings stay inside the firmware, and the lab's “hottest sensor” is a peak held since the card startedFound on 4 Oct. The host gets the integer mean of 34 minion-shire sensors plus peak-holds kept since the card's service processor started. Our dashboard shows that peak as the “hottest sensor”: aifoundry1's card 1 has shown 63 °C since 2 Oct at a 53–61 °C mean, and card 0's “123 °C hot spot” at a 65 °C idle was a peak left from September (C21). No command, trace or debug path gives the sensors one by one.all four cardscorrupts resultsNew workaroundNekko team CF14; us the dashboard's labels
C31A rejected read-only management query logs “Critical, SP Runtime Error”, although the card is fineFound on 4 Oct, reading the service processor's source against the kernel log. Every line the SP logs at error level counts as a runtime error, and with the threshold at 0 each one reaches the host as a Critical event. aifoundry2's first five such events (20, 22 and 25 Sep) each fell in the same second as a batch of our read-only queries, one of which the firmware rejected; the card ran kernels after each. Only the sixth, on 28 Sep, came with a hang (C27).all four cardswastes timeNewNekko team CF3
C32A card that falls off the PCIe bus still looks present, and a warm reboot does not bring it backaifoundry2's card dropped off the bus at 09:42 on 1 Oct and at 07:42 and 12:02 on 2 Oct, and has been off since. Each time its /dev nodes stayed, the driver stayed bound, and et-who called the card free; the kernel log showed only queue errors (SQ[0] sync: head mismatched, head_remote: -1), never “card lost”. Only sysfs tells: the card's link speed reads Unknown and its width 63 (still on 4 Oct). A plain reboot at 10:47 on 2 Oct left the slot empty; reboots that cut the slot's power brought the card back.aifoundry2; any cardwastes timeNew workaroundNekko team CF5; us H34
C33Nothing reads the cards' DRAM temperature, and DRAM refresh never speeds up on a hot cardFound in the source and in one measurement. The memory set-up leaves the controllers' temperature derating commented out (“not needed for bring-up”), nothing reads the LPDDR4X's own temperature register (MR4), and the memory shires have no sensor. In E53 (28 Sep) refresh stayed at its programmed rate at die means of 52–75 °C. LPDDR4X needs faster refresh above 85 °C, and the dies have run at 90–138 °C. No wrong result was seen up to an 81 °C mean; nothing hotter was checked.all four cardscorrupts results latentNewNekko team CF15
The host machines and access
H1Tailscale SSH asks for a browser check, and the check link can 404Unchanged. The check-mode steps and the 404 fix are in our repository's docs/lab-access.md and in Appendix A, but not yet on the public New user page that newcomers now follow (ours to add, AS3).all threeblocks workOpen workaroundNekko team AS3, DI1
H2Any account on one lab machine can log in as root on the othersUnchanged on 1 Oct. Since 2 Oct the lab's own onboarding uses the shared root login on purpose: each newcomer logs in as root once to create their own account (the lab's message to newcomers, and our New user page), so any narrowing must keep a way to create accounts (AS1, AS5).all threesecurityOpenNekko team AS1, AS5
H3aifoundry1's disk is full of user data: 95% of a single 452 GB poolFixed on 30 Sep at 22:25 PDT, at the owner's word: we deleted a departed user's public model checkpoints, the 27 files of 1 GB or more (118.1 GB of public models such as Qwen, Llama, Gemma, SmolVLM, RWKV, LFM and TinyLlama, which can be downloaded again). The account no longer existed, and nothing used the files (no process had them open, and no system setting or scheduled job named them); the smaller checkpoints, all code and a 10 GB compiled bundle were kept. /home went from 99% used (7.3 GB free) to 72% (116 GB free), and the pool from 95% to 71% (130 GB free of 452 GB). The other owners' data is unchanged (MO1); with no quotas the pool can fill again (PO4). Still so on 4 Oct: 115 GB free on /home (72%), and the pool 71% full with 128 GB free.aifoundry1blocks workFixed 30 Sep updatedNekko team PO4 (so it does not fill again), MO1 (the rest, no longer urgent); us U18
H4aifoundry1: ZFS permanent errors in four files, on a single disk with no backupsUnchanged: the 13 Sep scrub's 42 errors, no scrub since, no snapshot and no backup. The automatic scrub runs at 00:24 on Sunday 11 Oct. aifoundry1 was most likely switched off without a clean shutdown at about 12:55 on 2 Oct for the fan work (its journal stops mid-stream), another hard stop for this single disk.aifoundry1corrupts resultsOpenNekko team MO2, PO4; us U17
H5aifoundry1's 24 September kernel update stopped half-wayHolds: dpkg --audit is empty, and unattended upgrades ran cleanly on 26 and 27 Sep.aifoundry1blocks workFixed 25 SepUs U18
H6The reboot into kernel 7.0.0-34 has been pending since 24 Sep; done only on aifoundry3All three hosts run 7.0.0-34: aifoundry3 since 25 Sep, and aifoundry1 and aifoundry2 since Roman power-cycled the lab at about 15:07 on 30 Sep, after our link test hung aifoundry1 (C28). Section 4.6's checks passed on all three at 15:18–15:20.aifoundry1, aifoundry2wastes timeFixed 30 SepNekko team PO3 (a maintenance window for the next one)
H7Most unclean resets hit all three machines at once: their power is cut togetherNo unplanned cut since 18 Sep. All three restarted together, without shutdown records, at 15:07 on 30 Sep: Roman's deliberate power cycle after C28. aifoundry1 was most likely switched off without a clean shutdown at about 12:55 on 2 Oct for the fan work. A single machine should be powered off by itself, after sudo poweroff (SH2).all threecorrupts resultsOpenNekko team SH2
H8The boot never finishes: the splash screen waits forever on headless machinesFixed: all three hosts boot to the text target (U27), and after the power cycle of 30 Sep at about 15:07 all three reported running with no queued jobs (15:18–15:20).all threecosmeticFixed 30 Sep workaroundnothing left
H9aifoundry1's journal kept only about a dayaifoundry1's journal is at its 1 GB cap again and has grown by about 170–185 MB a day since 2 Oct, mostly sudo lines from our own live monitor, one a second (H35). It still keeps about ten days, about five at this rate. aifoundry2's is at 1.14 GB of its 2 GB cap.aifoundry1wastes timeFixed 25 Sepus H35 (stop the per-second sudo); aifoundry2's cap (root)
H10Time came from one NTP server over Wi-FiHolds: chrony is in sync with 8 sources on all three (offsets 0.09–0.33 ms).all threecorrupts resultsFixed 25 Sepnothing left
H11No Ethernet: all three machines run on Wi-FiThe mitigation holds (Wi-Fi power saving off); still no cable, and aifoundry2's roaming got worse: it switched access points 45 times on 30 Sep by 14:47 (3–25 a day on 26–29 Sep).all threewastes timePartly fixedNekko team SH4
H12CI runners run as root and can take a card at any timeUnchanged: both runners run as root, enabled and idle (no job since 15 Jun and 24 Jul), still polling GitHub.aifoundry1, aifoundry2corrupts resultsOpenNekko team MO3, PO2
H13aifoundry3's demo web app can launch jobs on the card as root, outside any lockUnchanged: the demo and the chatbot still run, with no requests since 25 Sep; disabling them no longer affects the driver (C2).aifoundry3corrupts resultsOpenNekko team MO3, AS1
H14numpy, venv, scipy and build packages were missing; aifoundry1 could not make a venv at allHolds: numpy 1.26.4, scipy and venv on all three.aifoundry1, aifoundry3blocks workFixed 25 Sepnothing left
H15Users could not read the kernel log, where the driver explains refused opens and card errorsHolds. (Users still cannot read the system journal: by design.)all threewastes timeFixed 25 SepNekko team DI1
H16Crashing programs left no core dumpsProven in use. Side effect: apport's coredump hook also wrote .crash copies (C19) until it was switched off, on aifoundry1 on 28 Sep and on aifoundry2 and aifoundry3 on 30 Sep (U20); aifoundry2's failed hook unit cleared with its reboot; only apport's crash report of 28 Sep is left there (C19).all threewastes timeFixed 25 SepNekko team DI1
H17Host CPUs ran in power-saving mode, so host-side timing jitteredHolds: the performance profile on all three.all threecorrupts resultsFixed 25 Sepnothing left
H18Months of pending updates, and a plain apt upgrade over Tailscale SSH kills itselfHolds; recurs by design without a maintenance window. Docker 29 / containerd 2 are held on aifoundry1 only, for the CI owners (aifoundry2 and aifoundry3 already run Docker 29.1.3 and containerd 2.2.1). On 4 Oct no security update was pending on any host; 28, 21 and 16 others were, among them kernel 7.0.0-38 (PO3).all threewastes timeFixed 25 SepNekko team PO3
H19Sudo, passwords and OpenSSH do not match the documented policySudo grew on 2 Oct: 7 members on aifoundry1, 18 on aifoundry2 and 3 on aifoundry3 (4 Oct). One new account got sudo on all three hosts that day, and a new file appeared in aifoundry2's /etc/sudoers.d; the New user page promises accounts with no sudo. Password logins stay off on aifoundry1 and aifoundry2. Our public accounts runbook still says to give sudo by setting a starting password (U7).aifoundry1, aifoundry2securityPartly fixedNekko team AS2; us U7, U22
H20Host firmware: 2021 BIOS versions, and SSD firmware with a known health bugUnchanged: BIOS F5, F5 and F6, and SSD firmware 3B2QGXA7 on all three. Since the 30 Sep boot only aifoundry1 has the IRQ 9 storm (aifoundry2's interrupt rate now matches aifoundry3's, so the storm comes and goes), and all three log the same ACPI BIOS error at boot (\ADBG, AE_ALREADY_EXISTS).all threecorrupts results latentOpenNekko team SH3
H21No console or out-of-band access, and the boot menu was hiddenThe 5 s boot menu holds; still nobody at a console, and on 30 Sep it mattered: aifoundry1 went down and nothing could restart it remotely (C28, SH8).all threewastes timePartly fixed updatedNekko team SH6, SH8
H22/tmp is wiped at boot, and coding agents keep their working files thereOur working files now live in the home directory (our rule since 30 Sep), so there is no /tmp copy to make, and U1 is closed. The rule itself is by design.aifoundry2 (all three)corrupts resultsPartly fixed updatedNekko team DI1
H23On every lab machine, ssh to another lab machine goes over the LAN, not TailscaleFixed on all three: each machine reaches the other two at their tailnet addresses (U24); aifoundry3's wrong line was corrected at 15:36 on 30 Sep.all threewastes timeFixed 30 Sep updatednothing left
H24aifoundry3 has no sshd: Tailscale is the only way inUnchanged: no openssh-server on aifoundry3. Installing it key-only is ours now (U23, waiting for the owner's decision).aifoundry3blocks work latentOpenUs U23
H25aifoundry2's tmux is a third-party snap that the Ubuntu package would breakUnchanged: the snap tmux stays held. Replacing it after the campaign is ours now (U21).aifoundry2wastes timeOpen workaroundUs U21
H26Background load on the measurement hosts, some of it oursIt grew again on 2 Oct, and the largest part is ours: our live monitor runs on all three hosts and, every second, checks the holders with sudo and reads each card's temperature through its management node (H35); the dashboard (every 10 minutes) and the History page (every 5 minutes) run from aifoundry2's crontab with probes into the other two hosts, with a watchdog and a node watcher every minute. Pausing all of it during a campaign is ours (U14).all threecosmeticOpenUs H35, U13, U14, U21
H27An idle login blocks anyone who follows the "nobody else logged in" etiquetteIt recurs with the new users: on 4 Oct two newcomers' logins had been idle since 2 Oct (on aifoundry1 and aifoundry3), and two of ours on aifoundry1. Coding agents now run in tmux by design, so only the card lock can be the rule (PO2, H36).aifoundry1, aifoundry3wastes timeOpen workaroundNekko team PO2; us close our idle sessions
H28aifoundry2's card never cools below the governor's 65 °C threshold, so its DVFS is almost never seenOut of service. The card has been off the PCIe bus since 12:02:01 on 2 Oct (its link reads Unknown, width 63), and the host has stayed on since 10:53 that day with the card in it. Its own temperature cannot be read; the host's drive, network-chip and CPU sensors have tracked aifoundry3's within about 2–3 °C since about 12:30 that day, so there is no sign that the card still heats. There is no workaround: only the cooling fix (SH5), then a power-cutting reboot to bring the card back (C32).aifoundry2corrupts resultsNewNekko team SH5 (the next visit)
H29aifoundry3's host copies memory at about half the other hosts' rate: it runs on one memory channelCause found on 27 Sep at 22:18 and confirmed as root on 28 Sep (dmidecode): one 32 GB DDR4-2666 DIMM, in ChannelA-DIMM1, three slots empty; a second DIMM is on-site work.aifoundry3corrupts resultsNew workaroundNekko team SH7
H30et-who prints a sentence when nobody holds a card and always exits 0, so a script took “free” for “held”Fixed on 28 Sep: et-who --check (0 free, 1 held, 2 failed) and the new idle sentence are installed on all three hosts; plain et-who still prints the same holder lines and exits 0.all threewastes timeFixed 28 Sepnothing left
H31Other accounts now work on aifoundry3, and one used its card; the report assumed only we didSeveral users at once is now normal: three people were onboarded on 2 Oct, and since then two other accounts have used aifoundry3's card and one aifoundry1's card 1 (the card-usage log). The newcomers took the card lock every time; one established account did not (PO2).aifoundry3corrupts resultsNew workaroundNekko team PO2, DI1
H32aifoundry3's journal has reached its 2 GB cap and will start deleting the July boot historyFixed on 28 Sep, and at risk again: since 2 Oct the journal grows by about 140–225 MB a day, about ten times the rate of 28 Sep, mostly sudo lines from our own live monitor (H35). At 2.6 GB of the 4 GB cap it fills around 11–13 Oct and then deletes the July boots again (the export of 28 Sep survives).aifoundry3wastes timeFixed 28 Sepus H35
H33Desktop services on the headless hosts fill the error logThe desktop part holds on all three: bluetooth, cups-browsed and the updater are off, and all three have booted headless since. The error counts have not been re-read (that needs the adm group). The new noise is ours: our live monitor's core dumps and its per-second sudo lines (H35).all threecosmeticPartly fixed workaroundUs H35; re-read the error counts (root or adm)
H34et-who --check covers the whole host and cannot see a card that is downet-who --check exits 1 if anyone holds any card node or lock on the host. On aifoundry1, where two people may now work at once, one per card, a script that gates on it waits for the other person's card. It lists holders only, so a card that fell off the bus keeps its /dev nodes and shows as free (C32). Each call runs sudo (H35). The banners and the onboarding brief say to read only your own card's lines and to check the link in sysfs.aifoundry1, aifoundry2wastes timeNew workaroundus et-who --card N with a link check (installing it needs root)
H35Our live monitor reads every card once a second without the card's lock, crashes now and then, and logs a sudo line every secondSince 16:05 on 2 Oct our live collector, a user service on all three hosts, reads each card's temperature once a second with a copy of ettelem, but only while et-who --check finds no holder on the host. Each read holds the single-opener management node for about 4 ms without the card's lock: about 177,000 opens of each of aifoundry1's cards by 4 Oct, and a tool started in that window gets “busy”. The reader has crashed in the vendor's libDM.so six times on aifoundry1 and twice on aifoundry3 (no reading lost), the failure C10 warns of. And each tick runs sudo, so aifoundry1 and aifoundry3 log about 3,600 sudo lines an hour, and their journals grow by 150–225 MB a day (H9, H32).all threewastes timeNewus take the card lock around each read, read the holders without sudo, report the crash (CF5)
H36The “is anyone else here” checks miss people, and nothing holds a card between runswho reads utmp, which has no entry for a Tailscale SSH command without a terminal: on 28 Sep it listed nobody on aifoundry3 while uptime counted three users. loginctl misses coding agents in tmux under linger: on 30 Sep aifoundry2 showed no sessions while two users had processes running, and every account made since 2 Oct has linger. The card lock lasts one run, so a newcomer setting up or between runs holds nothing.all threewastes timeNew workaroundNekko team PO2; us our repository docs
H37The root installs staged on 2 October were never made: two login banners contradict the cards' stateaifoundry1's login banner (30 Sep) still says card 0 overheats and must not be used, though its fan was replaced on 2 Oct and the dashboard and the New user page offer it. aifoundry2's does not say its card is out of service, and still invites newcomers to et-lab-start and the card lock. The installed et-lab-start (2 Oct 14:33, all three) hands the agent over at step 4 of the brief, which skips the simulator check. aifoundry1's card-usage logger skipped card 0 until 4 Oct; it logs it again. Corrected copies have been staged since 2 Oct and need one root session (U28).all threewastes timeNewus U28 (root)
Developing and measuring
D1The card's meters are coarse and filtered, and half of an idle card's power is on no meterNumbers replaced by E41, and the filter by E58 (29 Sep, pre-registered): a first-order average of 1.01–1.06 s on aifoundry3's rails and 1.08 s on card 1's minion and NoC rails, but 0.54 s on card 1's SRAM rail, so the cards' meters differ; aifoundry3's slower pass has a candidate cause (C23).all four cardscorrupts resultsOpen updated workaroundNekko team CF11, DI3
D2Some on-chip traffic starves the card's own meter on aifoundry2Unchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.aifoundry2corrupts resultsOpen workaroundNekko team CF11
D3Power follows die temperature, heat carries over between runs, and the room's airflow changesaifoundry2's chassis keeps its card hot enough to hide its governor (H28).all four cardscorrupts resultsOpen updated workaroundNekko team SH5
D4Common instructions trap in user mode: divide, square root, sine, 64-bit integer-to-float, double, the cycle CSRUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.all four cardswastes timeOpen workaroundNekko team RT4, CF11, DI3
D5Scratchpad addressing traps: offset 0 faulted once, and a global atomic through self ID 0x7F is a bus errorUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.all four cardswastes timeOpen workaroundNekko team DI3, CF11
D6The L1 data cache is not coherent and writes back whole 64 B linesUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.all four cardscorrupts resultsOpen workaroundNekko team RT3
D7One hot line stops a shire: hammering a global atomic stalls its home shire's memory path, with no errorUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.all four cardsblocks workOpen workaroundNekko team RT3, DI3
D8VPU register-file erratum 1.29: the compiler inserts no workaround and the simulator does not model itUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.all four cardscorrupts resultsOpenNekko team RT4, DI3, CD1
D9sys_emu and silicon disagree in ways a new user does not expectUnchanged: sys_emu and the cards are as on 25 Sep.sys_emuwastes timeOpen workaroundNekko team RT6
D10Current gp-sdk does not build against the lab's /opt/et, and there is no shared installUnchanged: still no /opt/gp-sdk on any host; it waits for one /opt/et (C18).all threeblocks workOpen workaroundNekko team RT2, RT6
D11Kernel ELFs built on different hosts hash differently, but the code is identicalUnchanged (cosmetic): the toolchain in /opt/et is as on 25 Sep.all threecosmeticOpen workaroundNekko team RT2
D12A stray write from a kernel leaves no trace on the card or the hostUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.all four cardscorrupts resultsOpen workaroundNekko team CF11
D13Cycle counting traps: hpmcounter3 reads 128 short, the cycle CSR traps, and evict_va is asynchronousUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.all four cardscorrupts resultsOpen workaroundNekko team RT3, DI3
D14Code on the lab hosts goes stale, or breaks, when it is updated in placeOne more instance, fixed and written down; the lab part is three lines in the onboarding page.all threewastes timeOpen workaroundNekko team DI1
D15Long or remote jobs over Tailscale SSH die, kill themselves, or expose their command linesHit again on 27 Sep; the trap in our checklist was fixed that night (D22). Command lines are still readable by every user on all three hosts (30 Sep).all threewastes timeOpen workaroundNekko team AS2, DI1
D16Our scripts/deploy-lab.sh used macOS tar flags and failed on LinuxHolds: scripts/deploy-lab.sh uses the portable tar wrapper.our repositoryblocks workFixed 25 Sepnothing left
D17A coding agent told not to touch the cards ran a real measurement block on a shared cardSafeguards extended for the heat code; fail-closed dry runs still to do.one lab hostcorrupts resultsOpen workaroundUs U13; Nekko team PO2
D18Our published notes and the lab-access page still carry the old diagnosesThe old diagnoses are corrected in the repository and the mirrored pages (25–27 Sep); the public accounts page's correction is prepared but not published; new errors moved to D23.documentationwastes timePartly fixedUs U7
D19Two host-to-card copies on one stream move less than one, and the link gives no full-duplex gainNarrowed on 29 Sep by E55 (pre-registered, three cards): only two copies on one stream collide, moving 0.49 of one, while one copy on each of two streams moves 1.01, so the loss is per stream; a shared DMA read engine and the IOMMU are refuted, and the IOMMU-passthrough boot (U26) is no longer needed. Both directions at once give 1.07–1.08× (E50). Every card and root port run MaxPayload 256 B and MaxReadReq 128 B (28 Sep).three cardswastes timeNew workaroundNekko team DI5
D20A small copy or an empty kernel costs hundreds of microseconds, mostly the runtime polling for completionE50: 556–566 µs for a waited empty kernel against 104 µs queued; 377–411 µs for a lone 4 KB copy.all threewastes timeNew workaroundNekko team RT5, RT2
D21Documents the open drop cites but does not contain, so parts of the chip must be inferredThe chip diagram and the memory-level pages mark these parts “inferred”.documentationwastes timeNew workaroundNekko team DI6, CD1–CD6
D22Our own checklist tells agents to run pgrep -af queue.sh over ssh, which always matches itselfFixed on 27 Sep at 23:34 and holding: AGENT.md §11 step 4 uses pgrep -af '[q]ueue.sh' and says why (commit a745199, pushed that night), and our card-behaviour notes and public newcomer brief carry the bracketed form. The 27 Sep update counted it as not committed yet.our docswastes timeFixed 27 Sepnothing left
D23Our docs carry new wrong or sensitive statements: card 1 “firmware DVFS”, a root SSH route, a stale numpy notePartly corrected: card 1's rows in AGENT.md and our card-behaviour notes are right, but the lib.sh comment is not. Its correction was reverted on 28 Sep to keep lib.sh at the bytes our locked experiments use, so it still says aifoundry3 has no system numpy, and it waits for the owner's decision on that lock. add-lab-user.sh still names the root route; that is the owner's call (U14).our repositorywastes timeNewUs U13, U14
D24Our PCIe probe held a card's lock for 12.8–14.9 s per run, over the lab's 10 s ruleFixed on 28 Sep: run_pcie.sh releases the lock between sub-tests, and its first card run (aifoundry1's card 1, 23:56) took at most 1.84 s per sub-test, with no lock wait; the stale copies on aifoundry1 and aifoundry3 were replaced.three cardscosmeticFixed 28 Sepnothing left
D25Reading the host CPU's energy needs root, so CPU-against-card energy comparisons rest on an assumed CPU powerThe host CPU's energy counter (/sys/class/powercap/intel-rapl:0/energy_uj) is readable by root only on all three hosts (checked 4 Oct), the kernel's default since a 2020 side-channel fix (CVE-2020-8694). So our 29 Sep comparison of CPU and card energy had to assume 125–251 W for the CPU. A policy choice, not a fault: the lab decides (MO8).all threewastes timeNewNekko team MO8; then us (root)

2. Requests for Roman#

The owner has root on the three machines, so everything that needs only root there, and no one else's decision, has moved out of this list into our own (section 2.9: on 4 October, 7 done, 2 partly done and 4 waiting for the owner's decision; one, the Gen3 test, took aifoundry1 down, C28), and these requests keep only what needs the Nekko team or the lab: the firmware, the driver and runtime upstream, the toolchain and simulator, the site and hardware, the tailnet, other people's data, accounts and services, documents only they have, and policies that bind every lab user.

Root could also carry out some requests that stay here, such as clearing other people's files, disabling the CI runners, locking accounts or changing aifoundry3's clock pin (MO1, MO3, AS2, CF2), but those are other people's data, services and accounts, so the decision stays with them and the lab. What the Nekko team would have to do, from the 27 September re-check, brought up to date on 30 September (SH8 and CF13 are new), with section 2.7, the chip diagram's asks, added on 1 October, and on 4 October (CF14, CF15, MO8, AS5 and SH9 are new). Each request says what to do, why, the problems it closes, the effort and who; the problems link back to their deep dives. 54 requests in 8 groups: Cards, driver and firmware (15) · Runtime and toolchain (6) · Machines and operations (4) · Access and security (4) · Site and hardware (9) · Documents and information to share (6) · The interactive chip diagram (6) · Policies (4). Everything else we do ourselves (sections 2.9 and 4.4). This list is error-specific: the firmware changes the measurements need (per-shire temperatures, input-side power, faster unfiltered power, ECC sources, a counter-event syscall) are rungs 11–16 of the hub's ladder, and the documents we asked for are its rungs 21–30; requests here point to them instead of repeating them. Per-shire temperatures are now also CF14, because the lab's own tools mislabel the peak the cards send (C30).

First, at the next visit
  1. SH5: the BIOS fan settings of aifoundry2, copied from aifoundry3: its card heated whenever the host idled, dropped off the bus on 1 Oct and twice on 2 Oct, and has been off it since 12:02 on 2 Oct (H28). The 2 Oct visit fixed aifoundry1's card 0 (SH1).
Then the five that matter most
  1. CF1 + PO1: give the cards real thermal protection. Then pick one firmware release that has the two upstream fixes.
  2. AS1 + AS2 + AS5: root between the machines and the demo's open port, who keeps sudo, and a way to make accounts without the shared root. Password SSH is off on both since 30 Sep (U22).
  3. MO1 + MO2: aifoundry1's disk. It is no longer full (116 GB free since 30 Sep, when a departed user's public model checkpoints were deleted at the owner's word), but it has four corrupt files in other users' data and a toolchain tree, and it has no backup.
  4. SH2: find out what cuts the power to all three machines at once, and fit a UPS.
  5. PO2: make the card lock the rule for everyone. Other accounts now work on aifoundry3 (one used its card on 27 Sep), and the CI runners and the demo still run as root with no lock.

Order. Within each group, by importance: how much the problem blocks or endangers work or data if nothing is done. Ties go to the cheaper item, then to the one that affects more machines or users. Effort is a rough estimate for someone who knows the system. The dot after each problem shows its status (green fixed, blue partly, amber open, violet new).

2.1 Cards, driver and firmware#

CF1 Real thermal protection on every card#

1 · EndangersEffort the answer, an hour; the firmware, days plus a signed imageWho firmware

What
In firmware, act on the hottest sensor (or on the mean plus a margin), enable the PVT controllers' over-temperature interrupts as a real die trip, assign max_temp, and report thermal and PMIC events to the host. For now, a one-hour answer: which of the lab's releases (1.2.0, 1.3.1, 1.4.1) carry upstream e024210bc (the safe state moves the PLL), 478275330 (the power sum no longer wraps above 65.535 W) and 7c6049087 (a failed voltage set no longer loops until the watchdog resets the card). By our mapping (C8), 1.2.0 and 1.3.1 predate all three, so the question is really about 1.4.1, which has no public source (CF9). Also say which sensor the PMIC's 75 °C alarm reads and whether it can fire, and what die temperature the chip is rated for: the preliminary datasheet leaves its ratings to a later release.
Why
Our reading of the governor source at the cards' releases found no die hard trip and no hot-spot response. The only thermal input is the integer mean of 34 sensors, and aifoundry2's card behaves so: in DV2's validation (28–29 Sep, E51) its clock held 800 MHz for a second or more after the hottest sensor read 67–69 °C, in 17 runs. DM_CMD_GET_MAX_TEMPERATURE returns 0. On 0.21.x (card 0) the governor does nothing while the card is idle, and 0.21.0 steps the clock up above 65.5 W. Card 0 drew 66–71 W at 115–117 °C on 25 Sep.
Since
4 Oct: no answer yet. et-platform is still untagged, its head still 836a4ab of 17 Jul. On 2 Oct aifoundry2's idle card heated from 45 to 138 °C and 21.9 to 133.8 W with nothing on the card acting (H28). The “hottest sensor” the host sees is a peak held since the card started, not the current hottest (C30, CF14).
Problems
C22 C21 C19

CF2 A clock pin on aifoundry3 that leaves the governor alive#

1 · EndangersEffort the pin, an hour; the firmware, daysWho aifoundry3's admin; firmware

What
Replace the 0 W TDP pin with a maximum-frequency or VMIN-table cap. In firmware, let power_throttling return when the operating point cannot go lower. Confirm once with an SP trace at INFO right after a reboot. The guard's reset gap (its marker survives a card reset) would be closed by our et-reset (U16), which would remove the marker and restart the guard if aifoundry3's admin agrees; the guard itself stays with that admin.
Why
At TDP 0 the SP's power task cannot leave the 600 MHz floor. After the first 65 °C crossing, the card makes no thermal step until the SP reboots. Every trace since 25 Sep is empty, and runs there reach 88–90 °C with nothing stepping in. Our runs stop themselves at 90 °C; other users' runs do not. It is rated Endangers because no card has a die trip (C22): the governor is the only thermal response, and on aifoundry3 the pin removes it, so a failed fan or a hotter workload would meet nothing but the 75 W board limit and a PMIC alarm that may not fire (CF1).
Problems
C23 C5 C6

CF3 Why aifoundry2's Master Minion hung, and which reset recovers a hung card#

2 · BlocksEffort the driver answer, hours; the firmware question, daysWho firmware, driver, platform

What
In firmware, find why aifoundry2's Master Minion stopped taking work at 02:50:53 PDT on 28 Sep (C27). Our hypothesis was a kernel launched while the governor's idle reset was taking the clock from 800 to 600 MHz; a later launch on the same card met the same step and ran normally (below), so that timing alone does not explain the hang. The runtime-error count is answered from the source: every error-level line the SP logs counts, with a threshold of 0, so the five earlier events were most likely our own rejected queries (C31). Say whether the SP's own 10 s Master Minion watchdog fired that night: in the source it only logs “MM Hung” and counts the hang, and nothing resets the Master Minion, although a narrower command for that exists (DM_CMD_MM_RESET); say whether that command is a supported recovery that leaves the rest of the card alone. In the driver, say why the sysfs per-card reset (soc_reset/reinitiate) re-attached the card but left the Master Minion hung, while the management reset (DM_CMD_RESET_ETSOC) recovered it, and which of the two is the supported recovery. Add a per-reservation clock pin tied to the card lock. The reset itself is done (U15), and a checked et-reset wrapper is ours (U16): installed on all three hosts on 30 Sep (aifoundry2's at 15:36).
Why
A hung card cannot run a kernel for anyone until it is reset. aifoundry2's card was down from 02:50 to 08:32 on 28 Sep, which held up the owner's DVFS validation, and the first reset tried, the sysfs one at 06:39, did not bring it back: launches at 06:41–06:47 still failed with Couldn't use the HPSQ. The management reset at 08:32 did, and a test kernel ran at 08:33. Without a reset that works, a hung card costs a whole-host power cycle, and if that cycles a supply shared by all three machines (H7), it resets everyone. The hypothesis rests on one case, so it is not established: the program started 0.6 s after the previous 7 s kernel ended, its 14 ms calibration kernel ran within a few tens of milliseconds of the clock change, and the next kernel never ran (board power stayed at the 26 W idle). The next two programs failed in the runtime's constructor while aborting the stuck command. Four seconds after the launch the kernel log reported SP Runtime Error, Runtime Error Count Beyond Threshold: 6; the five earlier such events on this card since 18 Sep left it running kernels. Of the night's 46 launches on this card it was the only one that started at 800 MHz and the only one met by a step down; 5 launches met by a step up within 0.5 s ran normally. The card has not hung since: in DV2's validation on the restored card (28 Sep 20:45 to 29 Sep 16:57 PDT, E51) all 52 launches returned 0, and one of them, at 23:34:18 on 28 Sep, was launched 0.50 s after the previous kernel ended and met the step to 600 MHz 0.16 s later (the hung launch met it at 0.13 s), then ran its 7 s normally. The launch records, the programs' output and the telemetry (every 100 ms) are ours to share.
Problems
C27 C14 C6 C7

CF4 Package the ET driver#

2 · BlocksEffort 1–2 daysWho platform

What
One versioned DKMS .deb built from a tagged et-platform commit. It should ship the modules-load.d entry, check after install that the module version is not empty, and replace every hand-built module.
Why
The 25 Sep hand fixes hold (C1-C3), but the next kernel or a hand copy can bring back the empty version string that locked aifoundry1's two cards out from about 18 to 25 Sep.
Problems
C1 C2 C3

CF5 A management path that survives a killed tool and names the holder#

2 · BlocksEffort daysWho driver, deviceLayer, libDM

What
Flush (or tag and drop) stale replies when a node is closed. Return a retryable busy error that names the holding PID, in the EBUSY kernel message or in sysfs. dev_mngt_service and et-powertop should finish the request on SIGTERM and catch exceptions instead of aborting. Errors should print what they read. Fix the udev rule (it matches nothing), and fix or document GET_FIRMWARE_BOOT_STATUS, GET_MAX_TEMPERATURE (returns 0) and low_power (on 0.21+ it only means board power ≤ 30 W).
Why
One kill -9 leaves a card unusable for the next opener until someone drains it. A busy node makes dev_mngt_service abort with a core (10 of them on aifoundry1 on 26 Sep).
Since
4 Oct: since 2 Oct our own temperature reader, a copy of ettelem built on the vendor's libDM.so, has crashed with a segmentation fault in that library six times on aifoundry1 and twice on aifoundry3 (and ettelem itself once), with no reading lost (H35). The udev rule still matches nothing.
Problems
C10 C11 C19 C1

CF6 Card temperature and power that anyone can read without taking the card#

2 · BlocksEffort daysWho driver

What
Have the driver expose die temperature and board power read-only through hwmon or sysfs, from the SP's periodic stats, with many readers allowed.
Why
Today a card's temperature can be read only through its single-opener management node. An overheating card cannot be watched without taking it from its user, a health check cannot warn, and monitors collide with tools: our card-0 guard held that node, without card 0's lock, during card 1's heat runs (27–28 Sep). Since 2 Oct the lab's live monitor reads every card's temperature once a second through the management node (about 4 ms, only while et-who --check finds no holder, and without the card's lock), so a management tool started in that window gets “busy” (H35).
Problems
C25 C21 C11 C26

CF7 Every stock tool addresses one card#

2 · BlocksEffort daysWho platform

What
Upstream ET_DEVICES into the stock deviceLayer, or make it honour -n. Rebuild dev_mngt_service and et-powertop with it. Ask the fork's author (account rehan) to upstream the change.
Why
On aifoundry1, a stock tool run for card 1 also opens card 0's management node. Using one card takes the other, so our telemetry passes had to lock the whole host, and on 26 Sep two of our own tools collided on card 0.
Problems
C13

CF8 aifoundry1 card 1: is DVFS off on purpose?#

3 · CorruptsEffort a 15-minute run on card 1, then the firmware team's answerWho the lab

What
No command reads the card's active-power-management flag: the management API has only the set (DM_CMD_SET_MODULE_ACTIVE_POWER_MANAGEMENT), which changes the card for everyone, and 1.2.0's closest public source sets it on at every boot. So: run one short cool-start test on card 1 since the 30 Sep and 2 Oct cold boots (busy below 65 °C, logging the clock, under both of aifoundry1's card locks, at the owner's call) to see whether it now steps. If it still never steps, the firmware team says why 1.2.0 never steps on this card; if that is by design, record it on the per-card sheet, or reflash under PO1.
Why
In every one of 359,657 samples the card read 600 MHz, including 11,446 below 65 °C (where 1.2.0 should step up) and readings up to 88 °C (where it should step down). It read 600 MHz again in all 10,033 samples of our 29 Sep runs, 1,170 of them busy and cool. A trace probe on 27 Sep saw no governor line at all.
Problems
C24 C7

CF9 Say which source each firmware release came from#

3 · CorruptsEffort hoursWho firmware

What
Tag every release in et-platform, return the commit in the firmware-revision command, and publish the source (or at least the commits) of 0.21.1 and 0.21.2.
Why
Users read December 2025 source for mid-2024 firmware. We mapped 1.2.0 to da192816a and 1.3.1 to cafe03fc3^; card 0's 1.4.1 (0.21.2) has no public source (the closest public commit is 50310b06b). The mapping, with each build's governor, has been in our repository since 28 Sep (docs/findings/14-card-behaviour.md); on 4 Oct the public et-platform still had no tags or releases, and its head was still 836a4ab of 17 Jul.
Problems
C8

CF10 Split the card commands by privilege#

3 · CorruptsEffort daysWho driver, firmware

What
Keep read-only telemetry open to everyone. Reset, trace level, the telemetry stats reset, temperature thresholds, active power management, TDP, frequency and rail voltage go to a group or to sudo helpers, and each state change is logged with the caller's UID.
Why
Any user can change what every other user measures. One reader's stats reset restarts every reader's rail averages (E41). Any user can also, silently, set the threshold to any byte or turn DVFS off, which removes a card's only thermal response. Setting a rail's voltage (DM_CMD_SET_MODULE_VOLTAGE) is open to every user too, and by our reading of the source for E54 (28 Sep; untested, since no voltage has been written) it is hazardous on the lab's builds: on card 1 (BL2 0.18.0) each NoC set rewrites the flash sector the boot voltage comes from, and on 0.20.0 (aifoundry2, aifoundry3) a failed or out-of-range set retries until the service processor's watchdog resets the card; the upstream fix, 7c6049087 (10 Oct 2024), is in none of the lab's builds. Until then, our banners list the other commands (U5, installed on all three hosts on 28 Sep), not yet the voltage set.
Problems
C12

CF11 Firmware telemetry and trap reports#

3 · CorruptsEffort days per item, plus a signed imageWho firmware

What
per-reading timestamps; unfiltered and input-side power; per-shire temperatures (the stock firmware has no per-shire readout); traps reported to the host with cause, PC and instruction; user-mode stores kept out of the PU and Maxion windows; the ECC sources. Also find which traffic starves the SP's management loop on aifoundry2 (D2). The details are the hub's rungs 11–16. The case for unfiltered power grew on 29 Sep: E58 measured the rails' averaging time at 1.01–1.06 s on aifoundry3 and 1.08 s on card 1's minion and NoC rails, but 0.54 s on card 1's SRAM rail, so filtered readings from different cards differ until each is undone with its own time constant.
Since
4 Oct: the per-shire temperatures in this list are now a request of their own, with how to get what exists today (CF14).
Problems
D1 D2 D4 D5 D12

CF13 Fix the driver's counter race upstream#

3 · CorruptsEffort the report, minutes; the fix, an hour plus a driver releaseWho driver

What
Report the race in et-platform and track it until the fix is in et-driver and in the lab's driver package (CF4); Filed on 4 October, publicly: et-platform issue #136 (the chips are end-of-life, so there was no case for a private report); a patch is ready for the maintainers. The fix is a reordering in the driver's queue set-up and teardown: show the counter files only after the queue tables exist, and remove them before the tables are freed. Until the lab runs a fixed driver, avoid reading the queue counters (mgmt_vq_stats, ops_vq_stats) during a card reset: a tool that reads them on a timer should pause while a card is being reset or its driver reloaded, and nothing should read them in a tight loop.
Why
A read at the wrong moment returns garbage counts, or, while the driver loads, crashes the reading program and can leave that card's next reset hung until a reboot (C29). Anyone can read these files, and lab tools, ours among them, read them on a timer to see whether a card is busy.
Problems
C29

CF14 Per-sensor temperature readings#

3 · CorruptsEffort the answer, an hour; the command, a day plus a signed imageWho firmware

What
A management command (a free number; 74 in the open source) that returns every temperature sensor the service processor reads: 35 entries, one per minion shire (0–33) and one for the I/O shire, each with where it sits, its current reading, and its hardware high and low since their last reset. Send the raw 12-bit code or millidegrees, not the whole degrees that pvt_ts_conversion() truncates to, plus a fault flag and the SP's time of the sample. The firmware already reads each sensor on its own (pvt_get_min_shire_ts_sample() and pvt_get_ioshire_ts_sample(), pvt_controller.c:645-665 at et-platform 353f20e), and pvt_get_and_print(…, PVT_PRINT_MINSHIRE_ALL, …) (:1573) fills all 34 shires but nothing calls it. Also write the same 35 readings into the SP statistics trace once per SP pass, as a custom event beside TRACE_CUSTOM_TYPE_SP_OP_STATS (dm_task.c:301-318), so a run gets a per-sensor series without holding the single-opener node. If a new command is too much for now, the smallest change is one call: print the per-shire temperature line the driver already has (pvt_controller.c:1183-1213) from the SP pass, the way the per-shire voltages are printed at DEBUG, which we already read from the trace (E6). In the existing temperature reply, add the current hottest sensor and its number, and a way for one reader to restart its own peaks without resetting everyone's. Document there that minshire_high and minshire_low are peak-holds kept since the SP started or its statistics were last reset, and that pmic_sys repeats the mean. Publish the sensor map, which needs three answers: confirm that sensor n sits in minion shire n (the firmware and the open RTL read it that way, while the firmware's voltage map scrambles the shires, pvt_controller.c:154-290); say where each shire sits on the die (the PRM's Figure 1-3 draws one sensor per tile, and our placement comes from mesh latency); and say whether the fused per-sensor calibration (35 sensors in the PRM's eFuse map) should replace RUN_1's nominal constants. For now, a one-hour answer: confirm that 1.2.0, 1.3.1 and 1.4.1 fill these fields the way the open source does (C8; the PVT driver is the same in 0.18.0, 0.20.0, 0.21.0 and 353f20e).
Why
The cards report one die temperature: the integer mean of the 34 minion-shire sensors, each truncated to whole degrees before averaging (pvt_controller.c:1303-1340, thermal_pwr_mgmt.c:732-772). The two extremes sent with it are hardware peak-holds kept since the card's service processor started. They are not the hottest sensor now, but our dashboard has labelled them “hottest sensor”: aifoundry1's card 1 has shown 63 °C there since 2 Oct while its mean idles at 53–61 °C, and card 0's readings of 25 Sep, 28 Sep and 2 Oct showed 123 °C at a 64–65 °C idle, a peak left from its September overheating (C30). Read in 1 s windows, the hottest sensor runs 1–4 °C above the mean, idle or loaded (E53). In aifoundry2's runaway it reached 144 °C at a 138 °C mean; the gap was 1–2 °C below 80 °C, 2–4 °C at 100–119 °C and 6 °C at the end (H28). The governor, and any trip built on the mean, lets the hottest shire run that much further (CF1, C22), and nobody can say which shire it is. The 35 sensors sample continuously at 12 bits, but no command, trace record, debug path or sysfs entry carries them one by one (C25). CF11 lists per-shire temperatures among its telemetry items; this request is that item on its own, with the code that does it.
Until then
How to get per-sensor (disaggregated) temperatures today, from the most to the least available:
  • From the host, without the card: nothing per sensor. The lab dashboard's Live section and the History page show each card's mean once a second, and the peak since the card started, which they label “hottest” (C30).
  • Holding the card: ettelem sample gives temp_c.minshire [mean, lowest, highest], where lowest and highest are peaks held since the service processor started or its statistics were last reset, and temp_c.ioshire [current, low, high], the I/O shire's single sensor, the only one read on its own. Nothing else names a shire.
  • The hottest sensor over a short window: only by resetting the statistics at the start of the window (ettelem sample --reset-ms) and reading the peak at its end. The reset is global: it restarts every reader's peaks and the SP's power statistics, and as our tool sends it also turns off the SP's statistics trace (C12). So the onboarding brief forbids it, and only the owner's registered runs use it, holding the card's lock (E41, E53). It still gives the hottest value, not which shire.
  • Which shire: not at all on the lab's firmware. The firmware already reads each sensor; one print call, or a command, would export them (above). Running changed firmware needs a signed image (CF12, C9), and the per-shire readout is also a rung of the hub's ladder. The simulator cannot stand in: sys_emu's sensors return one fixed value (RT6).
Problems
C30 C22 H28 C21 C25 C12

CF15 The DRAM's temperature, and the chip's thermal limits#

3 · CorruptsEffort the answers, an hour; the firmware, days plus a signed imageWho firmware, hardware

What
  1. Name the DRAM part on the cards and its temperature grade.
  2. Say whether the memory's temperature derating was left off on purpose. In firmware, read the LPDDR4X's temperature register (MR4) and enable derating, or at least report the DRAM's temperature to the host.
  3. Give the die's rated junction temperature, the temperature its timing was signed off at, and its design life (CF1 asks only for the rated temperature).
  4. Say what the PMIC's “system temperature” measures, and whether the PMIC can power-cycle the board.
  5. For the ~120 °C episode the lab reported earlier: which card and firmware, what stopped, and whether a reset was needed.
Why
Hot cards may lose DRAM data silently: refresh stays at its cool-card rate, and nothing reads the memory's temperature (C33). The dies have run at 90–138 °C (C22, H28). Our overheating study (E53, 28 Sep) could not answer these from the repository.
Problems
C33 C22 C21

CF12 A way to run changed firmware#

4 · Wastes timeEffort a decisionWho Nekko and the vendor

What
Sign community builds on request, or keep one lab card on a development key.
Why
C9 blocks one kind of work, running changed firmware, for everyone. It is ranked 4, not 2, because no card or machine is unusable and nobody's current work depends on it: the specific firmware changes the measurements need are asked for directly in CF11 (the hub's rungs 11–16), which is the practical route for now.
Problems
C9

2.2 Runtime and toolchain#

RT1 Fix the log-level race in the runtime#

3 · CorruptsEffort hoursWho runtime

What
Register the VLOG_* levels before any thread starts (in rt::IRuntime::create, or with a static initializer), or make g3log's level lookup non-inserting. Build the reference runtime once with ThreadSanitizer.
Why
Without our one-line workaround, host programs crash about once in 100 launches, 1.08 s in. On aifoundry3's runtime that was 9 SIGSEGV and 3 SIGABRT on 25–26 Sep, and it lost campaign measurements. The race is in every build, and another account now runs programs on aifoundry3.
Problems
C17

RT2 One reference /opt/et on every host, with a matching gp-sdk#

3 · CorruptsEffort 1–2 daysWho the lab, with account rehan and aifoundry3's admin

What
Build one tagged commit at one build type (RelWithDebInfo) and install the same bytes everywhere, with /opt/gp-sdk built against it. Put experimental builds in /opt/et-<name>.
Why
The three hosts run three different runtimes. E50 (27 Sep) shows that launch and small-copy latencies follow each host's build, so host-side results cannot be compared across machines. No host has a gp-sdk that builds.
Problems
C18 D10 D11 D20

RT3 Library defaults that do not trap new users#

3 · CorruptsEffort daysWho et-common-libs, gp-sdk

What
Allocate per-hart outputs line-aligned by default (D6); ship a barrier that does not spin on a global atomic, and audit the sync helpers (D7: one hot line stops its shire's memory path, with no error); ship a corrected counter read (D13); and state each hart's user-mode stack size, 4,160 B (KERNEL_UMODE_STACK_SIZE in et-common-libs' layout.h, with hart 1's stack directly below hart 0's), or put a guard between them: on 29 Sep a kernel of ours whose hart 0 used 4,240 B was caught only by sys_emu's memory checker, as a write hazard on hart 1's stack (E59).
Problems
D6 D7 D13

RT4 Toolchain: reject what traps, and handle erratum 1.29#

3 · CorruptsEffort daysWho toolchain

What
A CPU option that rejects user-mode-trapping instructions at compile time (divide, square root, sine, 64-bit int-to-float, double, the cycle CSR); the erratum 1.29 workaround, or at least an assembler warning (on 29 Sep the compiler's own save and restore of the callee-saved f registers, around a function whose asm clobbers them, produced a type-A hazard that sys_emu -vpurf_warn caught before any card run; E59); make -vpurf_check skip firmware PCs so it can run by default.
Problems
D4 D8

RT5 Cheaper completion in the runtime#

4 · Wastes timeEffort daysWho runtime

What
interrupt-driven or short adaptive polling for completion, and document the two polling constants.
Why
An empty kernel costs 556–566 µs when waited for, against 104 µs when queued. A lone 4 KB copy costs 377–411 µs, and back-to-back round trips lock onto one of two polling modes (E50).
Problems
D20

RT6 sys_emu and gp-sdk#

4 · Wastes timeEffort daysWho sys_emu, gp-sdk

What
Upstream the firmware-preload fix: in upstream main (836a4ab of 17 July, still the head on 4 Oct), GenericLauncher still gives sys_emu no firmware to preload. --emit-relocs has been upstream since 2 May (7e3b4c2), later than the lab's gp-sdk 06605ab, so a current gp-sdk has it (RT2). Publish a list of what sys_emu does not model (among them: its PVT controller is a stub, and every temperature sensor and voltage channel returns one fixed value, about 51 °C and 402 mV, so no thermal or DVFS behaviour can be tested there; CF14), add a TensorSend partner checker, and make the VPURF checker see tensor writes to the f registers (erratum 1.29 type F), which it ignores (E59, 29 Sep).
Problems
D9 D10

2.3 Machines and operations#

MO1 Free aifoundry1's disk#

1 · EndangersEffort the owners' hoursWho the account owners, the lab

What
Mostly done on 30 Sep: at 22:25 PDT, at the owner's word, we deleted a departed user's public model checkpoints (the 27 files of 1 GB or more, 118.1 GB), which no account or process used any more. What is left is no longer urgent: the owners may still delete or move what they do not need (account rehan, about 188 GB; account roman, about 54 GB), and the lab may clear what remains in the two root-owned homes, /root/et-jobs-deploy, and the two unused Docker images and three stopped containers. Our part is done: our 25 Sep snapshots were destroyed on 28 Sep at 08:33 (U3), which gave back about 2.4 GB.
Why
The pool was 95% full, with 7.18 GB available on 30 Sep at 14:38 PDT, down from 7.98 GB just after U3 on 28 Sep (5.0 GB on 27 Sep, 6.4 GB just after the 25 Sep cleanup), and the rest of it was user data. After the deletion /home is 72% used (116 GB free, from 99% and 7.3 GB) and the pool 71% (130 GB free of 452 GB; fragmentation 25%, from 59%). A full pool stopped the 24 Sep kernel update half-way and squeezed the journal; with no quotas (PO4) it can fill again.
Problems
H3

MO2 aifoundry1's corrupt files, and a backup#

1 · EndangersEffort hoursWho the owners, the lab

What
The owners restore or delete the corrupt files: account rehan's two .gguf files, and one git object in another user's home (re-clone it). The lab deletes the unused /usr/src/riscv-gnu-toolchain/build-* trees (4.9 GB, which hold the corrupt cc1plus), or tells us they are unused. Then decide on a backup: at least zfs send of rpool/USERDATA to another machine. The scrub, zpool status -v and zpool clear that follow are ours (U17).
Why
42 errors since the 13 Sep scrub, on a single disk with no backups that has had 11,276 unsafe shutdowns. Since our snapshots went (28 Sep, 08:33), zpool status -v lists these four files and nothing else. On 30 Sep it still reported the 13 Sep scrub's 42 errors and no scrub since. The next automatic scrub runs on Sunday 11 Oct at 00:24, so the files are best dealt with before then.
Problems
H4

MO3 The CI runners and aifoundry3's demo#

3 · CorruptsEffort 30 minutes eachWho the CI owners, aifoundry3's admin

What
Disable the runners if the hackathon is over, or reinstall them unprivileged with every job holding the lock of the card it uses. Retire aifoundry3's dead runner (its unit and /root/actions-runner-hf-hackathon). Disable the demo and chatbot, or have them take the card lock (their network exposure is in AS1). chown root:root the chatbot's and llama-server's unit files, which are owned by a user account (the demo's own unit file is root's).
Why
All of these run as root and can take a card at any moment, outside every lock. The runners have been idle since 15 Jun (aifoundry1) and 24 Jul (aifoundry2) but still poll GitHub. The demo had no requests from 25 to 27 Sep (its journal, read as root on 27 Sep). On 30 Sep the demo and chatbot were still enabled and running, and the runners on aifoundry1 and aifoundry2 were running with no job.
Problems
H12 H13

MO8 Let users read the host CPU's energy, or say no#

4 · Wastes timeEffort a decision; then a rule, minutes as rootWho Roman

What
Decide whether users may read the host CPU's energy counter. Either keep it root-only (comparisons stay assumed, and the reports say so); give a group read access with a udev or tmpfiles rule, accepting the side channel; or let a small root service publish package energy coarsely, once a second, which removes it. We make the change as root once the lab decides.
Why
CPU-against-card energy comparisons now rest on an assumed CPU power (D25).
Problems
D25

2.4 Access and security#

AS1 Root between the lab machines#

1 · EndangersEffort 30 minutes in the tailnet policy, 10 for the demoWho the tailnet admin, aifoundry3's admin

What
Decide whether lab machines need root on each other. If they do not, restrict root to an admin group from their own devices, with check mode, and limit machine-to-machine rules to autogroup:nonroot. Write down the root paths that remain. Close the one root web service on the tailnet as well: aifoundry3's demo runs as root and answers on its tailnet address with no login; bind it to localhost, put a login in front of it, or limit that port in the tailnet policy.
Why
Any account on one lab machine can open a root shell on the others with no password or check. We used it, at the owner's request, for the 25 Sep fixes, the 27 Sep read-only re-check and the root steps of 28 Sep. It was still open on 1 Oct. Since 2 Oct the lab's onboarding uses the shared root login on purpose (AS5), so any narrowing must keep a way to create accounts. Closing it does not have to end the owner's root: an admin group, from their own devices and with check mode, is enough for the steps of section 2.9.
Problems
H2 H13

AS2 Accounts, sudo and the adm group#

1 · EndangersEffort an hourWho Roman

What
Decide who keeps sudo and the NOPASSWD rules on aifoundry1 and aifoundry2, from the figures we already have (aifoundry2 had 17 members on 30 Sep, 16 with usable passwords, and each of the two hosts one NOPASSWD rule; our read-only review, U22, can refresh them at any time). Lock dormant accounts. Archive the homes that have no account (two on aifoundry1, one on aifoundry2). Keep the adm group, which can read every log, to admins. Password logins in sshd are now off on both hosts (U22: aifoundry1 since 28 Sep 20:51, aifoundry2 since 30 Sep 14:37).
Why
Both hosts accepted password logins on port 22 from the LAN until U22; since 30 Sep both offer only public-key logins. Who holds sudo has grown: 18 members on aifoundry2, 7 on aifoundry1 and 3 on aifoundry3 (4 Oct, against 17 and 6 on 30 Sep). On 2 Oct one new account got sudo on all three hosts, and a new file appeared in aifoundry2's /etc/sudoers.d; the New user page promises accounts with no sudo, so the lab should confirm the policy. The accounts are other people's, so the lab decides who keeps what. Our own accounts runbook suggested giving sudo by setting a password; our correction is prepared (U7) and is published once the owner agrees.
Problems
H19 D15

AS5 Accounts without the shared root, and one identity per person#

1 · EndangersEffort an hour to decide, then the tailnet policyWho Roman, the tailnet admin

What
Since 2 Oct newcomers start from the shared administrator login and create their own account there; that is the lab's rule, sent to newcomers on 2 Oct, and our New user page follows it. Decide who keeps root afterwards; give a way to create an account that does not hand out root (an admin, or a small helper that creates the account and does nothing else); make each login map to one person. Any narrowing of root under AS1 must keep a way to create accounts.
Why
Work done as root cannot be traced to a person, and everyone onboarded this way has had root on the machine. A related gap in the tailnet's login rules is described in a private note for Roman, not on this page.
Problems
H2 H19

AS3 Tailscale: check mode, and the invite#

2 · BlocksEffort an hourWho the tailnet admin

What
Consider a longer check period (or a device-posture rule) for the member-to-lab-machine rule, so an overnight job does not need a browser. Put the check-mode steps and the 404 fix in the invite. They are in Appendix A and in our repository's docs/lab-access.md, but not on our public page New user? Start here, which says only that each ssh may print a login link to approve; adding the 404 fix there is ours. A way into aifoundry3 when Tailscale is down, OpenSSH key-only, is ours (U23).
Problems
H1

2.5 Site and hardware#

SH1 aifoundry1 card 0: cooling first, then a reseat#

1 · EndangersEffort an hour on siteWho on site

What
Done, 2 October. On site, card 0's fan was found broken and replaced. Our acceptance test that afternoon, run remotely: 49 °C idle against 65 °C before, and 52–53 °C (peak 56 °C) under 8 minutes of sgemm bursts, cooler than card 1 (59–60 °C, peak 63 °C) under the same test (C21); the “123 °C hot spot” quoted before was a peak held since September (C30). The card is back in service. Left: the installed banner still warns against card 0; the corrected one is staged and needs one root install (U28). The link's corrected errors came back at about 7 an hour after the 2 Oct boot, far below the old flood; we watch them (C15).
Why
The card reached 98–102 °C in short test runs and 115–117 °C just after. The driver counted 18 board-power overshoots (75.0–75.75 W against 75 W) and 10 thermal-throttle events on 25 Sep, unchanged on 30 Sep. The firmware has no die trip (CF1). Its PCIe link, trained at 16 GT/s x8, still logs about one corrected error a second while idle (918,587 at 23:20 on 27 Sep; 1,147,012 at 14:41 on 30 Sep), so the errors do not come from load. The flood also leaves users only about ten hours of aifoundry1's kernel log.
Problems
C21 C15

SH2 Power#

1 · EndangersEffort half a day, plus the purchaseWho on site

What
Find what cuts the power to all three machines at once: building outages, a breaker, a shared strip or PDU, or someone cycling it. Put the machines on a UPS with NUT, and keep a log of on-site power work.
Why
On aifoundry1, 11,276 of 12,089 power cycles were unsafe shutdowns; the cuts cluster on 17 and 23 Jul across machines, and the last unplanned one was on 18 Sep. On 30 Sep at 15:07 all three restarted together again, without shutdown records, when Roman power-cycled the lab on purpose, and aifoundry1 was most likely switched off without a clean shutdown at about 12:55 on 2 Oct for the fan work. Ask that a single machine is powered off by itself, after sudo poweroff. (aifoundry1's fall on 30 Sep was not a power cut: it followed our link test, C28.) Single-disk hosts with no backup (MO2) take every cut.
Problems
H7

SH8 Power-cycle aifoundry1, and read why it went down#

2 · BlocksEffort ten minutes on siteWho on site; then us, remotely

What
Done. Roman power-cycled the lab at about 15:07 on 30 September, and aifoundry1 came back with both cards working (C28). We read the previous boot's journal as root on 1 October: the last entry is at 14:41:17, four seconds before the retrain, with no kernel error, no machine check and no panic record, so the host froze instantly. Card 0 is back in service since its fan was replaced on 2 Oct (SH1), and nobody should repeat the link test (U25).
Why
Our Gen3 test of card 0's link, run as root with card 0 idle and locked, retrained its root port to 8 GT/s at 14:41:21, and within twenty seconds the whole host stopped answering, on Tailscale and on the LAN. The lab has no console or out-of-band access (H21, SH6), and the hosts do not reboot after a panic, so nothing can bring it back remotely. Everything on it is down: the logged-in sessions, both cards and the CI runner.
Problems
C28 H21 C15

SH3 One console visit, in this order#

2 · BlocksEffort 2–3 hours at the consolesWho on site

What
Back up first (MO2). Enrol the driver's MOK (mokutil --import /var/lib/shim-signed/mok/MOK.der, then confirm at boot). Update the SSDs from 3B2QGXA7 to 5B2QGXA7. Bring all three BIOSes to the same current version. Afterwards, check Secure Boot and "power on after AC loss".
Why
The firmware of aifoundry1 and aifoundry3 is in Setup Mode, so a BIOS reset that restores the default keys could turn Secure Boot on and hide every card (C20). The 980 PRO firmware has a known health bug. The BIOSes date from 2021. aifoundry1 and aifoundry2, on BIOS F5, have had an ACPI interrupt storm (IRQ 9 disabled after 100,001 interrupts and polled; since the 30 Sep boot only on aifoundry1), and all three log the same ACPI BIOS error at boot.
Problems
C20 H20

SH4 Ethernet#

3 · CorruptsEffort 15 minutes and three cablesWho on site

What
Plug enp7s0 in on all three machines (the same 10GbE port on each). NetworkManager brings it up by itself, and Wi-Fi stays as the fallback.
Why
All access, root included, goes over Wi-Fi. aifoundry3 roams between two access points up to 16 times a day at about -70 dBm, and aifoundry2 switched access points 45 times on 30 Sep by 14:47 (3–25 a day on 26–29 Sep).
Problems
H11

SH5 aifoundry2's card cooling: the BIOS fan settings#

1 · EndangersEffort 30 minutes at two consolesWho on site

What
On site, at the BIOS of aifoundry3 and then aifoundry2:
  1. At aifoundry3 first (same board, a Gigabyte Z590 AORUS MASTER, same slot; its card idles at 54–57 °C): press Del at boot, F2 for Advanced Mode, then Smart Fan 5 (F6). Photograph the settings of every SYS_FAN header, and note each fan's speed in PC Health Status. These are the settings to copy.
  2. At aifoundry2, in Smart Fan 5, for every SYS_FAN header: Fan Stop → Disabled; Fan Speed Control → Manual with at least 50% at every temperature (or Full Speed); if Temperature Input offers a PCIe or system sensor, use it instead of CPU; Fan Fail Warning → Enabled.
  3. In PC Health Status, with the machine idle, write down each fan's RPM: a fan at 0 RPM is stopped or dead.
  4. Settings → Platform Power: AC BACK → Always On, so the machine starts again after a power cut (SH2); ErP → Disabled.
  5. Save with F10. Then look inside the case: does any fan blow across the card (it sits in the second CPU slot, with the top slot empty)? If none does, fit one aimed at the card.
  6. Check afterwards: on the lab dashboard, aifoundry2's card should stay below about 70 °C with the host idle. It reached 82–83 °C idle on the night of 30 September.
Why
On 1 October at 09:42 aifoundry2's card dropped off the PCIe bus while idle, and stayed off until a full-reset power cycle at 06:45 on 2 October brought it back (H28). Its temperature follows the host's CPU load, not its own work: 63–67 °C while the host was busy, 70 °C when it idled in the evening, 82–83 °C idle from 02:30 on 1 October (with idle power up from 29 to 37.5 W), and it failed minutes after the host went idle again. The chassis fans most likely follow the CPU temperature and slow down or stop when it idles; no fan speeds are visible from the host. aifoundry3's card, on the same board in the same slot, stayed at 54–57 °C all night. No card has a die-temperature cut-off (C22). 2 October: after each full-reset boot the idle card ran away and dropped off the bus within about 70 minutes, the second time watched once a second up to 138 °C and 134 W; on the same leakage curve as aifoundry3's card, its air is about 20 °C hotter (H28).
Since
4 Oct: not done at the 2 Oct visit. The card has been off the bus since 12:02 on 2 Oct, and the host has stayed on since 10:53 that day with the card in it, against this page's advice; the host's own sensors have tracked aifoundry3's within about 2–3 °C since, so there is no sign that the card still heats. Powering aifoundry2 off until the visit is the owner's call (the host also runs our dashboard and history jobs). The test after the fix: the idle card levels off below about 70 °C within 20 minutes; above 95 °C, switch off.
Problems
H28 D3 C7

SH6 Out-of-band access#

4 · Wastes timeEffort a purchase and an hourWho the lab

What
An IP-KVM (such as PiKVM) on at least aifoundry1, or a named on-site contact for every maintenance window. The GRUB menu is now shown for 5 s, but nobody is at a console to use it. On 30 Sep it mattered: aifoundry1 went down during our link test, and nothing could restart it remotely (C28, SH8).
Problems
H21 C28

SH7 A second memory DIMM for aifoundry3#

4 · Wastes timeEffort minutes on site, plus the DIMMWho on site

What
Add a matching 32 GB DDR4-2666 DIMM in memory channel B at the next visit, then we re-run the copy test.
Why
aifoundry3 has one DIMM, in ChannelA-DIMM1, with the other three slots empty, so it runs single-channel memory (read as a user on 27 Sep 22:18, U10). Its host memcpy runs at 9.2 GB/s, against 17.4 and 21.4 GB/s on the dual-channel hosts, so its staged copies to the card run at 5.2 GB/s against 7.2–7.8.
Problems
H29

SH9 Research help on site: a thermocouple, a spare card, and one low-clock run#

4 · Wastes timeEffort an hour on siteWho Roman; whoever goes on site; the card owner

What
  1. Permission to tape a thermocouple to a card's heatsink base during a visit: negligible risk, and it reads the heat path the die sensors cannot.
  2. If there is a dead or spare card, lend it for a delidded or IR-window view of the die; a bare lid shows almost nothing spatially.
  3. A go-ahead for our pre-registered 100 MHz run on one card, under its lock. It changes the card's clock, so it needs the lab's agreement (C12); 100 MHz is the lowest clock the firmware can set, and anything lower needs a signed bootloader (C9, CF12).
Why
Our heatsink and low-clock studies (30 Sep–1 Oct) ended with these asks. The die sensors give only a mean and a peak since boot (C30, CF14), so these are the remaining ways to see where heat sits on the die.
Problems
C30 D3

2.6 Documents and information to share#

DI1 Adopt our public onboarding page#

2 · BlocksEffort 30 minutesWho Roman

What
Review our public page New user? Start here (public since 30 Sep, linked from all three login banners since that evening; its brief for coding agents grew from Appendix A, U12) and link it from #community-lab. Since 2 Oct its first step follows the lab's rule: newcomers join the tailnet through Roman, then create their own account on one machine and work only as that account; confirm that step (AS5). It does not yet cover Tailscale check mode or the 404 fix (ours to add, AS3). It covers the card table, the 10 s rule and the 90 °C stop; et-who, et-who --check and the card lock; never kill -9 a card tool, and the one-line drain; setsid nohup for long jobs, killing by PID, and no secrets on ssh command lines; that /tmp is cleared at boot; dmesg and coredumpctl; the one-line crash workaround for host programs; the card commands that affect everyone; and that coding agents follow the same rules.
Problems
C10 C17 H1 H15 H16 H22 H31 D14 D15

DI2 Adopt the per-card sheet we drafted#

3 · CorruptsEffort 30 minutesWho Roman

What
Review our draft (Appendix B, U12) and keep it current after any reflash or policy change; our public New user? Start here brief carries the same card table, kept current. For each card it gives the firmware, the DVFS state, the TDP policy, the idle state, the boot clock, the thermal behaviour and the known quirks, plus aifoundry3's pin and the reason for it. Today three cards are in service and behave three ways: aifoundry3 is pinned and latched; aifoundry1's card 1 never steps; card 0 idles at 300 MHz and, since its fan was replaced on 2 Oct, runs cooler than card 1. aifoundry2's card is out of service: idle, it heats until it drops off the bus (H28, SH5). Appendix B below still describes 27 Sep.
Problems
C4 C5 C23 C24 H28

DI3 Errata and the programming guide#

3 · CorruptsEffort days of writingWho Nekko

What
Publish the list of user-mode-trapping instructions (D4); the reserved scratchpad regions and which address formats support atomics (D5); the hot-line rule with its two errata (D7); each card's silicon stepping, whether erratum 1.29 applies, and whether the compiler is meant to avoid it (D8; on 29 Sep its own register save and restore produced a hazard, E59); each hart's user-mode stack size (4,160 B, with no gap between the harts' stacks, E59); the counter's late carry and evict_va's asynchrony (D13); the meters' filter, refresh periods and rail coverage (D1; on 29 Sep we measured a first-order average of 1.01–1.08 s on most rails, but 0.54 s on card 1's SRAM rail, E58); and release notes for 1.2.0, 1.3.1 and 1.4.1, including what 1.4.1's low_power means.
Problems
C4 D1 D4 D5 D7 D8 D13

DI4 Two messages to pass on#

4 · Wastes timeEffort two messagesWho Roman

What
Tell aifoundry3's admin that our 25 Sep drop-in makes et-board-clock-guard load the driver first, that the guard trusts its boot marker after a card reset (C6), and that on 30 Sep at 14:39 the machine was changed as root (U27, U20, U24, U16, U18): it boots without the desktop from its next boot (the demo, the chatbot and the clock guard still start), bluetooth, cups-browsed and apport's core-dump hook are off, /etc/hosts names aifoundry1 by its tailnet address, and et-reset and a daily et-lab-health timer are installed. Since then aifoundry3 has also gained our card-usage logger (et-usaged, a root service, 30 Sep), et-lab-start (30 Sep, updated 2 Oct), new banner lines, our once-a-second live monitor, which reads the card's temperature (H35), and accounts that newcomers create themselves from the shared login (AS5); tell that admin these too. Ask account rehan to upstream ET_DEVICES (CF7).
Problems
C2 C6 C13

DI5 The PCIe DMA engine#

4 · Wastes timeEffort an emailWho Nekko

What
Say how the runtime and the Master Minion's DMA worker serve two copies queued on one stream, and whether one copy per stream is the intended way to overlap them. The engine's documentation, or the register values the firmware programs (channels, read-request size, element split; hub rung 29), would also explain the missing full-duplex gain. Our IOMMU-passthrough boot (U26) is no longer needed.
Why
Two host-to-card copies in flight on one stream move half as much as one (0.49 of one on all three cards; E50, and E55 on 29 Sep), but one copy on each of two streams loses nothing (1.01, E55), so the loss is per stream: E55 refuted a shared DMA read engine and the IOMMU. The link also shows no full-duplex gain (E50). On 28 Sep we read the PCIe side as root: every card and its root port run MaxPayload 256 B and MaxReadReq 128 B (a quarter of PCIe's 512 B default); whether that limits the rate of a single copy is untested (U10).
Problems
D19

DI6 The chip documents the open drop cites but does not include#

4 · Wastes timeEffort emails; some may need a licence decisionWho Nekko

What
The NoC manual and spec, the floorplan with the die orientation, the minion shire's floorplan, the Data Book or a final datasheet (the preliminary datasheet leaves its ratings and the package's thermal information to a future release, so no junction limit is published), the card's schematic and board file, the memory shire's description or its programmed DDR registers (on aifoundry2 two DRAM lines that differ only in PA[17] read as a row conflict, unlike on the other two cards; E57, 29 Sep), a way to read the NoC router and bridge counters, and the cited design documents (the hub's rungs 21–28). Also settle the storage-cell question: the Shire Cache Specification and the datasheet describe SRAM macros.
Problems
D21

2.7 The interactive chip diagram#

On 1 October the owner asked what the Nekko team could send to make the chip diagram and the memory levels more authentic. The documents behind both pages are asked in DI6 (the hub's rungs 21–28) and in the hub's rungs 37–42; these six ask for less, which may be easier to share: answers, an image and a check of the drawings. An answer can be public, or private to the diagram's author, who would publish only what the Nekko team approves; for another company's IP, such as TSMC's process or the PHYs, only what its terms let you state publicly (CD6).

CD1 Which of the open RTL's minion options the silicon has#

4 · Wastes timeEffort an email: a word or two for eachWho Nekko

What
Say which of the open RTL's options the ET-SoC-1's minion has: the hand-tuned (MMI) multiply-add and 64-bit adder, which the RTL can select but does not contain; the ET-custom vector register file, L1 data array and TenC buffer, which the drop has only as behavioural models; and the transcendental tables as synthesised logic, a ROM or a latch table. core-et's CONTRIBUTING.md says the RTL was taped out twice, at 1,088 cores and at 8: is the open Erbium branch's minion the ET-SoC-1's? The hub's rung 40 asks the same for the caches, rung 37 what MMI and the L1's cells are.
Why
The diagram draws the minion from this RTL, flags 18 facts "that the ET-SoC-1 silicon matches it block for block is not confirmed", and marks the vector register file unknown. The RTL's two register files differ in when a write commits (core-et-main's VPU README), and erratum 1.29 (D8) is a write-timing erratum in that register file, so the answer also says which of the two to simulate it with (our reading).
Problems
D21 D8

CD2 Does each minion shire have its own supply?#

4 · Wastes timeEffort an emailWho Nekko

What
Say whether the 34 minion shires' supply inputs stay separate in the package and on the die, joined only on the card, and whether any shire has its own regulation or power gating on the die. The voltage regions are in the Power Spec (the hub's rung 28).
Why
Esperanto's Hot Chips 33 talk says "Each Minion Shire has independent low voltage power supply inputs that can be finely adjusted to mitigate Vt variation effects and enable DVFS", but the dev card, by its manual, feeds every minion shire from one three-phase regulator output, and the diagram gives the minions one supply. The hub's rung 5 reads the per-shire voltages the firmware logs as IR drop on that one rail; separate inputs on the die would change that reading.
Problems
D21

CD3 A die image the page may show#

4 · Wastes timeEffort an email and a decisionWho Nekko

What
A photograph or an annotated plot of the die, with a scale bar and the side it is seen from (bumps or back), and permission to show it on the diagram's public page with the credit you want. If nothing else can be shared, permission to show the die plot of the Hot Chips 33 talk (slide 20) would do.
Why
The diagram shows no image of the silicon: its die is a grid of tiles drawn from that plot, which has no scale bar or stated viewing side and is the mirror image of the firmware's NoC map (the hub's rung 22). We have found no published photograph of the die.
Problems
D21

CD4 The die's numbers, if the floorplans cannot be shown#

4 · Wastes timeEffort the die and the tile, an email; the areas, an hour if the A0 reports are at handWho Nekko

What
If the floorplans of DI6 (the hub's rungs 22 and 23) cannot be shown, rounded numbers would do: the die's width and height, and the side the published die plot shows; a minion shire tile's width and height; where a shire's PLL sits; and, if the A0 reports are at hand, the area and cell count of a minion, a neighbourhood, a cache bank, the mesh stop, the I/O and PCIe shires' blocks, the vector unit, a lane and its multiply-add.
Why
The diagram scales the published die plot to 570 mm² and reads the rest from its pixels: a die of 25.6–25.8 by 22.1–22.2 mm and a 3.72 mm hop (3.64–3.74 over three readings), by which Heat per millimetre divides measured energies. Below the tile every size is an order-of-magnitude guess (a minion about 500 µm across, a lane 120 µm), and every transistor count is ours. Ranked 4, like DI6: the pages mark all of these inferred.
Problems
D21

CD5 Check the diagrams' inferred and unknown parts#

4 · Wastes timeEffort a few hours, with our checklistWho Nekko

What
Open the chip diagram and the memory levels as published on 1 October and, for each part drawn dashed (inferred on the chip diagram, unknown on the memory levels), say right, wrong or can't say; we would send them as a checklist of about 30 parts. A yes or no is enough: the documents that would settle each part are DI6 and the requests above.
Why
Of the chip diagram's 834 facts on 1 October, 161 are inferred and 14 unknown; the memory levels have 41 unknown. Experiments settled some on 28–29 Sep (E55–E57), but most can be settled only by someone who has seen the design, and to a reader who does not open a part's source, a wrong inferred part looks as solid as a measured one.
Problems
D21

CD6 What the vendors' terms let you state publicly#

5 · HygieneEffort an email, as far as the terms allowWho Nekko

What
Only what TSMC's and the IP vendors' terms let you state publicly: the number of metal layers, and which carry the power grid and the mesh links; and the outlines and sizes of the PCIe and LPDDR4X PHY macros. No private answer is asked here.
Why
The diagram's wiring stack gives no layer count ("TSMC does not publish N7's stack heights"), and it sizes the PHYs to an order of magnitude (the PCIe PHY about 2 mm, a DDR PHY about 1 mm). Ranked 5: each answer refines one drawing.
Problems
D21

2.8 Policies#

PO1 A firmware policy#

1 · EndangersEffort a decision, then a windowWho Roman

What
One release for all four cards that contains the two upstream fixes in CF1 and the voltage-set fix 7c6049087 (C12), reflashed in a window with a recovery path. Or keep the mix and publish the per-card sheet (DI2). Either way, state each card's active-power-management setting.
Why
The three releases differ in known thermal and power bugs as well as in behaviour. Card 0's release may step the clock up above 65.5 W, and it does nothing while idle. From the source (untested), card 1's 1.2.0 writes every NoC voltage set into the flash sector its boot voltage comes from, and on the two 1.3.1 cards a failed set loops until the watchdog resets the card.
Problems
C4 C21 C22 C24 C8 C12

PO2 The card lock is the rule for everyone#

2 · BlocksEffort a messageWho Roman

What
Replace the "nobody else logged in" etiquette with "a card is busy when et-who shows a holder or its lock is held". Every CI job, the demo, every user and every coding agent takes /run/lock/etsoc-shire<N>.lock for the card it uses. Announce it on Discord and in the banner; our public New user? Start here brief already teaches it (flock -n with a 10 s timeout), and the dashboard and card-usage log now show who holds each card and whether they took the lock.
Why
Another account used aifoundry3's card on 27 Sep 04:35–05:13, and a kernel-launch error followed. We do not know whether it took the lock, and the plans had assumed nobody else used that card. A second, different account logged in to aifoundry3 at 22:43 the same day. On aifoundry1, one idle login from 18 Sep blocked the old etiquette (and the reboot) for 12 days, until the host went down on 30 Sep (C28).
Since
4 Oct: the newcomers of 2 Oct took the lock for every program they ran (15, none longer than 5 s); one established account ran 22 programs on aifoundry3 without it between 1 and 2 Oct, 4 of them holding the card for 15–28 s. The banners still teach a blocking flock without -n or a timeout (ours to fix). aifoundry2's CI runner is still online although its card is out of service.
Problems
H31 H12 H27 C11 D17

PO3 A standing maintenance window#

3 · CorruptsEffort a decision, then an hour a weekWho Roman

What
Weekly, announced with wall and in the banner, so that every user can plan around it. In it we run, as root, the upgrades (through a detached unit, not a Tailscale SSH session; a full upgrade each time, since the unattended upgrades on all three hosts skip noble-updates), the reboots and the post-boot checklist. The CI owners test Docker 29 and containerd 2 before those are unheld on aifoundry1. On 4 Oct, 28 updates waited on aifoundry1 (6 of them the held Docker and containerd packages), 21 on aifoundry2 and 16 on aifoundry3, among them kernel 7.0.0-38 on all three, which the unattended upgrades skip; the ET driver is built only for 7.0.0-31 and -34, so that kernel needs the window and the post-boot checks.
Problems
H6 H18 C14

PO4 Storage#

3 · CorruptsEffort hoursWho Roman

What
per-user quotas (for example zfs set userquota@<user>=100G on the home dataset), one shared read-only /models dataset, zfs send backups, and an alert at 85% full (our daily health check, U18, would report it). On 30 Sep at 14:38 rpool was 95% full, with 7.18 GB available, and /home 99% used; at 22:25 a departed user's public model checkpoints were deleted at the owner's word, which left them at 71% and 72% (130 and 116 GB free), but nothing yet stops the pool from filling again.
Since
4 Oct: 115 GB free on /home (72%), the pool 71% full; still no quota on any home. et-lab-health warns at 85% every day, in the journal.
Problems
H3 H4

2.9 Ours now, with root#

The owner has root on aifoundry1, aifoundry2 and aifoundry3 (confirmed on 28 September at 08:30), so what needs only root there, and no one else's decision, is ours: it moved here from the requests above. On 28 September 2 items were done, and 4 ran on aifoundry1 that evening with the owner's approval. On 30 September at 14:37–14:41 PDT the owner ran our commands for the rest that wait on no one else: the headless target and the desktop services are now done on all three hosts (U27, U20); password logins are off on aifoundry2 as well (U22); the peer names are pinned on all three hosts (aifoundry3's wrong line corrected at 15:36) (U24); and et-reset and the daily health check are installed on all three (aifoundry2's at 15:36) (U16, U18). aifoundry3's admin should hear of these changes (DI4). Our Gen3 test of card 0's link (U25) then took aifoundry1 down (C28). So 7 items are done, 2 partly done and 4 wait for the owner's decision (the scrub, tmux and the timers, OpenSSH on aifoundry3, and since 4 October the installs we staged on 2 October, U28); U25 must not be repeated, and U26 is no longer needed (D19). Each item says what changed and how to roll it back. They are grouped by the kind of work, with the requests' group names, and sorted by importance within each group; runtime and toolchain, documents and policies have none, since those need the Nekko team or the lab (the IOMMU-passthrough boot and the headless target, which came from a document request and a policy, are machine work and sit under Machines and operations). “From” names the request each came from; the table below says where every request of 27 September went. Their state is also in section 4.4.

Cards, driver and firmware

U15 Reset aifoundry2's hung card#

2 · BlocksDoneFrom CF3 (its first step)Where aifoundry2

What
Recover the card whose Master Minion hung at 02:50:53 on 28 Sep (C27), without power-cycling the host.
Result
Done on 28 Sep at 08:32:45, with the owner's approval. The sysfs reset at 06:39:27 re-attached the card (the kernel log: enabling device, added peer-to-peer DMA memory) but left the Master Minion hung: our test launches at 06:41, 06:43 and 06:47 still failed (exit status 1; the two we saved, at 06:43 and 06:47, with Couldn't use the HPSQ. Perhaps the Master Minion is hanged?). The management reset, which needs no root (any user can send it, C12), recovered it: the kernel log shows Mgmt: Device is resetting, then the card re-enabled and its DMA memory re-added, and the tool reported success at 08:32:52. At 08:33 a 1 s test (our sparsity_host, the fma test on one minion) ran 3 launches, all ok at 600 MHz (0.599 GHz by its own count), and held the card for 1.65 s. Logs: labreport2/fix28/c27/, the commands and outputs in labreport2/fix28-aifoundry2.txt, and /root/labfix-20260928/ on aifoundry2.
What ran
# 06:39:27, as root: the per-card reset through sysfs (did not recover the Master Minion)
echo 1 > /sys/bus/pci/devices/0000:02:00.0/soc_reset/reinitiate
# 08:32:45, as our user, after et-who showed no holder: the management reset (recovered it)
dev_mngt_service -m DM_CMD_RESET_ETSOC -n 0
Risk to others
None beyond the card, which had run no kernel for anyone since 02:50. The host was not rebooted, so nothing else on aifoundry2 stopped and our /tmp files (H22) were not at risk.
Problems
C27 C14
U16 et-reset: a checked card reset#

2 · BlocksDone on all threeFrom CF3 (the et-reset wrapper), with CF2's guard gapWhere all three

What
A small command, et-reset N, that refuses while et-who --check shows any holder, holds the card's lock while it sends the management reset (the one that recovered C27; the sysfs reset did not), waits for the card to come back, and logs who ran it. On aifoundry1 a reset of card 1 also holds card 0's lock, because the stock dev_mngt_service -n 1 also opens card 0's management node (C13). Installed on all three on 30 Sep (Result below); the command has been in the repository's tools/lab/ since 28 Sep (commit 5538de2, described in tools/lab/README.md), dry-tested on all three hosts; its reset path has not run. It would be installed the way that README installs et-who. The version written leaves aifoundry3's clock guard alone and prints a reminder to tell that admin, so the guard part below waits for them. Who may reset a card is the lab's rule (PO2, CF10; today users ask the lab admin, C14, and Appendix A says not to reset a card yourself), so until the lab decides, only root and the machine's sudo group can run it. On aifoundry3 it would also remove the clock guard's boot marker and restart the guard, so that the 600 MHz pin comes back after a reset (C6: the gap that CF2 named), but only if aifoundry3's admin agrees (through DI4): the guard is that admin's service, which CF2 leaves with them.
Command
install -o root -g sudo -m 0750 tools/lab/et-reset /usr/local/bin/et-reset    # each host
# aifoundry3 only, once its admin agrees: let the wrapper restart the guard without a password prompt
printf '%s\n' '%sudo ALL=(root) NOPASSWD: /usr/bin/rm -f /run/et-board-clock-guard.ok, /usr/bin/systemctl restart et-board-clock-guard.service' \
  > /etc/sudoers.d/et-reset && chmod 0440 /etc/sudoers.d/et-reset && visudo -c
# what et-reset N does:
#   et-who --check || exit 1
#   flock -n /run/lock/etsoc-shire<N>.lock dev_mngt_service -m DM_CMD_RESET_ETSOC -n <N>
#   aifoundry1, card 1: flock -n /run/lock/etsoc-shire0.lock flock -n /run/lock/etsoc-shire1.lock dev_mngt_service ... -n 1
#   aifoundry3, if its admin agrees: sudo rm -f /run/et-board-clock-guard.ok && sudo systemctl restart et-board-clock-guard.service
#   logger -t et-reset "card <N> reset by uid $(id -u)"
Risk to others
A reset ends whatever runs on that card. The wrapper refuses while the card's lock (for aifoundry1's card 1, either card's) is held or a process has its node open, so only someone running without the lock could lose work. Installed for root and the sudo group only, it gives no one a power they lack: any user can already send this reset with no check at all (C12), and the sudoers line only spares the sudo group a password prompt for two commands they can already run with sudo. Opening it to every user is the lab's call (PO2, CF10). On aifoundry3 the guard part changes its admin's service, so it waits for that admin. Roll back: remove the two files.
Result
Installed on all three hosts: aifoundry1 and aifoundry3 by the owner on 30 Sep (14:39), aifoundry2 by us at 15:36 that day. The version installed never touches aifoundry3's clock guard (C6). Who may run it stays the lab's rule (PO2, CF10).
Not on aifoundry2 now
Do not use it on aifoundry2's card while it is out of service: a reset that brought it back would restart the idle runaway (H28, SH5); the card needs its cooling fixed and a power-cutting reboot (C32).
Problems
C14 C6

Machines and operations

U3 Destroy our 25 Sep ZFS snapshots on aifoundry1#

1 · EndangersDoneFrom MO1 (our part)Where aifoundry1

What
Our 34 snapshots of 25 Sep (@labfix-20260925-A and -B, taken before that day's changes and upgrade) held about 2.4 GB of a pool that was 95% full (H3).
Result
Done on 28 Sep at 08:33:40, with the owner's approval, after a dry run that listed only our snapshots (no holds, and no other snapshot on the host). zfs reported 832 MB, 1.31 GB, 105 MB and 171 MB reclaimed, and no labfix snapshot is left. Available space rose from 5.70 to 7.98 GB on rpool and from 1.36 to 1.65 GB on bpool. zpool status -v now lists only the four damaged files of H4 (labreport2/fix28/a1/s01-real.out).
What ran
zfs destroy -rv rpool/ROOT/ubuntu_jleyxm@labfix-20260925-A
zfs destroy -rv rpool/ROOT/ubuntu_jleyxm@labfix-20260925-B
zfs destroy -rv bpool/BOOT/ubuntu_jleyxm@labfix-20260925-A
zfs destroy -rv bpool/BOOT/ubuntu_jleyxm@labfix-20260925-B
Risk to others
None to other users: only our snapshots went. It removes the rollback path of the 25 Sep upgrade (snapshot B), which three days of holding changes no longer need.
Problems
H3 H4
U17 Scrub aifoundry1's pool, once the owners have dealt with their files#

1 · EndangersWaits for the owner's decisionFrom MO2 (the scrub)Where aifoundry1

What
When the owners have restored or deleted the four damaged files (MO2), scrub the pool, check that no new error appears, and clear the old ones. The next automatic scrub is due on 11 Oct.
Command
zpool scrub rpool          # pause with: zpool scrub -p rpool
zpool status -v rpool      # when it ends: the error count, and the files they are in
zpool clear rpool          # only once that list is empty
Risk to others
The scrub reads all 433 GB in use on the single disk, so while it runs (likely an hour or more) every user's disk access on aifoundry1 is slower: run it with the machine idle and no measurement running. It changes no data.
Result
Not run: on 30 Sep at 14:40:59 aifoundry1's pool showed no scrub since 13 Sep. The next automatic scrub runs on Sunday 11 Oct at 00:24.
Problems
H4
U18 Run et-lab-health every day#

4 · Wastes timeDone on all threeFrom MO4Where all three

What
The read-only health check that we installed on all three hosts on 28 Sep (U8), in a daily timer, so that its WARN lines reach the journal every day (and a mail or the banner, if the lab wants). It catches in minutes what took days to notice this month: an empty driver version, a module missing for a kernel, a PCIe error rate, the power and thermal event counters, a full pool, a half-done dpkg run, a stale guard marker. The two units are written in tools/lab/README.md (28 Sep, with et-lab-health rev 3, which is safe to run from a timer); both were installed on 30 Sep (Result below), and all three hosts run rev 3 (checked 4 Oct).
Command
printf '[Service]\nType=oneshot\nUser=nobody\nExecStart=/usr/local/bin/et-lab-health\nSuccessExitStatus=1\n' \
  > /etc/systemd/system/et-lab-health.service
printf '[Timer]\nOnCalendar=daily\nRandomizedDelaySec=1h\nPersistent=true\n[Install]\nWantedBy=timers.target\n' \
  > /etc/systemd/system/et-lab-health.timer
systemctl daemon-reload && systemctl enable --now et-lab-health.timer
journalctl -u et-lab-health        # its output
systemctl is-system-running        # still "running" on a day with WARN lines
Risk to others
Read-only: it opens no card node and changes nothing; it reads sysfs, /proc, the kernel log and the package, unit and pool state (dmesg, systemctl, dpkg, zpool) for a few seconds a day as nobody. The check exits 1 whenever it prints a WARN, which aifoundry1 does every day (card 0's events, and until 30 Sep the pool at 95%); SuccessExitStatus=1 keeps that from marking the unit failed, which would leave the system “degraded” and fail section 4.6's systemctl is-system-running check. On aifoundry3, which has its own admin, tell that admin first. Roll back: systemctl disable --now et-lab-health.timer, then remove the two files.
Result
Installed on all three hosts with et-lab-health rev 3: aifoundry1 and aifoundry3 by the owner on 30 Sep (14:39), aifoundry2 by us at 15:36 that day; first runs succeeded and the systems stay running.
Problems
C1 C3 C15 C16 C21 H3 H5
U19 Reboot aifoundry1 and aifoundry2 into 7.0.0-34#

4 · Wastes timeDone (by the power cycle)From MO5Where aifoundry1, aifoundry2

What
The reboot pending since 24–25 Sep. Both new driver modules and the initramfs are ready, and 7.0.0-34 is GRUB's default entry. It is also the first test of aifoundry1's boot-time driver load (C2). Set the headless target first (U27), copy our /tmp files on aifoundry2 again just before (U1), and run the checks of section 4.6 afterwards. The reboot is scheduled only if no one holds a card, and ten minutes ahead so that logged-in users are warned. Running it is ours, but on aifoundry1 it also needs the OK of the account logged in there and someone on site, so H6 stays with the lab (PO3).
Command
# aifoundry2 only, as our user, just before:
#   nice -n 10 ionice -c3 rsync -a /tmp/claude-1019/ ~/claude/private/tmp-claude-1019-20260927/
et-who --check && shutdown -r +10 "This machine reboots in 10 minutes into kernel 7.0.0-34"
# shutdown -c cancels it; afterwards: the checks of section 4.6, then et-lab-health
Risk to others
It ends every session and job on that host. On aifoundry1 that includes the account logged in since 18 Sep (H27), so it needs that user's OK, and someone reachable on site in case the ZFS root does not come back (no console, H21; SH6). On aifoundry2 it ends the owner's own session and wipes /tmp. The 5 s GRUB menu can still boot 7.0.0-31.
Result
Done, though not as planned: our link test hung aifoundry1 on 30 Sep (C28), and Roman power-cycled all three machines at about 15:07. All three now run 7.0.0-34, and section 4.6's checks passed at 15:18–15:20.
Problems
H6 C2
U28 Install what was staged on 2 October: the corrected banners, the current et-lab-start, the root-login notice and et-opens#

2 · BlocksWaits for the ownerFrom U5 (our banners), DI1Where all three

What
Install what we corrected and staged on 2 Oct (H37): aifoundry1's banner (card 0 back in service) and aifoundry2's (its card out of service); the current et-lab-start on all three; then the notice for root logins and the et-opens tracer, which names the short card opens the usage log records as unseen (both in tools/lab/README.md). Drop the card-0 exclusions from aifoundry1's et-usaged options, so card 0's queue activity is logged again (done 4 Oct). aifoundry3's changes go to its admin first (DI4).
Command
# as root, from the repository checkout; each copy is checked against the staged one first
# aifoundry1
echo 'e843bd9ff868fe2db7a4b5b34e330f87cfe1cacec83a07c8d96c070962bfe958  tools/lab/motd-aifoundry1' | sha256sum -c - && cp -p /etc/motd /etc/motd.bak-20261004 && install -m 0644 tools/lab/motd-aifoundry1 /etc/motd
# aifoundry2
echo '5a860aaf55359db9ebb1f268256a99de79261fd251693aa326f2b577b7bd6382  tools/lab/motd-aifoundry2' | sha256sum -c - && cp -p /etc/motd /etc/motd.bak-20261004 && install -m 0644 tools/lab/motd-aifoundry2 /etc/motd
# each host (aifoundry3 once its admin agrees)
echo 'b0551c7debed3bba1e6d3aa1d7e8670e8154ab4a4fbc94e477f07f0aef4fe247  tools/lab/et-lab-start' | sha256sum -c - && cp -p /usr/local/bin/et-lab-start /usr/local/bin/et-lab-start.bak-20261004 && install -m 0755 tools/lab/et-lab-start /usr/local/bin/et-lab-start
Risk to others
None to running work: these are text files and a script that only reads. Roll back: copy the .bak-20261004 files back.
Result
Partly done. On 4 Oct at 13:41 PDT the root-login notice and the et-opens tracer were installed on aifoundry1, after its probe check; et-opens has logged since, and the card-use section can now name every open there. Still waiting: the corrected banners on aifoundry1 and aifoundry2 (the installed ones date from 30 Sep), the current et-lab-start on all three (the installed one is the 2 Oct 14:33 version), the notice and the tracer on aifoundry2, and aifoundry3's share, which goes to its admin first. At 15:14 the card-0 exclusions were dropped from aifoundry1's et-usaged options, once the owner confirmed card 0 fixed; it logs card 0's use again, and our installer no longer demands them.
Problems
H37 C21 H28
U26 One boot with IOMMU passthrough, to test the copy rates#

4 · Wastes timeNo longer neededFrom DI5 (the IOMMU-passthrough boot, rung 35)Where aifoundry2

What
Two host-to-card copies at once move half as much as one (D19). One boot of aifoundry2 with iommu=pt (the hub's rung 35), with E50's copy tests run again, tells whether the IOMMU's address translation is part of it. Do it in the same window as the reboot (U19).
Command
echo 'GRUB_CMDLINE_LINUX_DEFAULT="$GRUB_CMDLINE_LINUX_DEFAULT iommu=pt"' \
  > /etc/default/grub.d/61-labfix-iommu-pt.cfg && update-grub
et-who --check && shutdown -r +10 "aifoundry2 reboots in 10 minutes (one boot with iommu=pt)"
# after the copy tests:
rm /etc/default/grub.d/61-labfix-iommu-pt.cfg && update-grub
et-who --check && shutdown -r +10 "aifoundry2 reboots in 10 minutes (back to the normal boot)"
Risk to others
Two reboots of aifoundry2, with the same cost as U19 (every session there ends, /tmp is wiped). For that one boot the devices on aifoundry2 do DMA without translation, which drops the IOMMU's protection against a faulty device; nothing else changes for other users.
Result
Not needed: on 29 Sep E55 (pre-registered, three cards) found that two copies collide only when both are on one stream (0.49 of one), while one copy on each of two streams loses nothing (1.01), and refuted the IOMMU and a shared DMA read engine as the cause (D19). What serves one stream's two copies in turn goes to DI5.
Problems
D19
U27 Headless boot on all three#

4 · Wastes timeDone (from each host's next boot)From PO5Where all three

What
Boot the three machines to the text target instead of the desktop, before their next reboot. Otherwise every boot waits on the splash screen until someone runs plymouth quit (H8). It also removes about 43 greeter processes per machine, and most of the bluetooth and notifier noise (H33).
Result
Done on aifoundry1 on 28 Sep at 20:51:12, with the owner's approval, and by the owner on 30 Sep on aifoundry2 (14:37:52) and aifoundry3 (14:39:50): systemctl get-default reads multi-user.target on all three (checked 14:41–14:48). All three have booted with it since (30 Sep and 2 Oct), and on 4 Oct all three read multi-user.target and running. Roll back: systemctl set-default graphical.target.
Command
systemctl set-default multi-user.target      # takes effect at the next boot
# roll back: systemctl set-default graphical.target
Risk to others
After the next boot there is no graphical login at a monitor, only a text console, so anyone who uses the desktop at the machine loses it. Nothing changes before that reboot. On aifoundry3, which has its own admin, tell that admin first.
Problems
H8 H6 H33
U20 Quiet the desktop services#

5 · HygieneDone on all three; one clean-up left on aifoundry2From MO6Where all three

What
Turn off bluetooth, cups-browsed and the firmware-updater snap's notifier on all three machines. On aifoundry1 also stop apport's duplicate .crash reports: they come from apport's hook on systemd-coredump, which keeps every core anyway (H16). Leave pam_lastlog alone: the line is optional, and a bad PAM edit breaks every login.
Result
Done on aifoundry1 on 28 Sep at 20:51, with the owner's approval, and by the owner on 30 Sep on aifoundry2 (14:37:54) and aifoundry3 (14:39:52): bluetooth and cups-browsed disabled and stopped, the firmware-updater snap disabled, and apport's drop-in on systemd-coredump linked to /dev/null on all three (checked at 14:41–14:48). systemd-coredump still keeps every core. Left on aifoundry2: apport's own crash report of 28 Sep (the failed unit cleared with the reboot of 30 Sep, and the host reads running on 4 Oct); removing it is rm -f /var/crash/_usr_share_apport_apport.0.crash, as root.
Command
systemctl disable --now bluetooth.service cups-browsed.service
snap disable firmware-updater
# aifoundry1 only: mask apport's drop-in on systemd-coredump
mkdir -p /etc/systemd/system/systemd-coredump@.service.d
ln -s /dev/null /etc/systemd/system/systemd-coredump@.service.d/apport-coredump-hook.conf
systemctl daemon-reload
Risk to others
Small on headless servers: a Bluetooth keyboard or a network printer used at the machine would stop working, and the desktop's firmware-update pop-up goes away (fwupdmgr still works). systemd-coredump still keeps every core. On aifoundry3, which has its own admin, tell that admin first. Roll back: systemctl enable --now the two services, snap enable firmware-updater, and remove the link, then daemon-reload.
Problems
H33 C19
U21 After the campaign: aifoundry2's tmux, and the timers#

5 · HygieneWaits for the owner's decisionFrom MO7Where aifoundry2; all three

What
Replace aifoundry2's snap tmux with Ubuntu's package once no tmux session runs there (H25), and stop the sysstat and plocate timers for the length of a measurement campaign (H26).
Command
# aifoundry2, only when no tmux runs for any user:
pgrep tmux || { snap remove tmux && apt-get install -y tmux; }
# any host, for the length of a campaign (systemctl start them afterwards);
# plocate-updatedb.timer only where it is installed:
systemctl stop sysstat-collect.timer sysstat-summary.timer plocate-updatedb.timer
Risk to others
Removing the snap would end any tmux session still running, which is why it waits until none runs. With the timers stopped, locate goes stale and sysstat's history has a gap for the campaign. On aifoundry3, which has its own admin, tell that admin first.
Problems
H25 H26

Access and security

U22 Turn off password logins in sshd, and review sudo#

1 · EndangersDone on aifoundry1 and aifoundry2; the sudo review waitsFrom AS2 (password logins, and the sudo review)Where aifoundry1, aifoundry2

What
On aifoundry1 and aifoundry2, OpenSSH accepts password logins on port 22 from the LAN (H19). A drop-in turns them off; Tailscale SSH and key logins are unaffected. It matches the lab's own model: docs/lab-access.md says logins go through Tailscale SSH, not sshd passwords, and accounts made by its script have no password. Separately, a read-only list of who has sudo and adm, and the NOPASSWD rules, refreshes the figures AS2 already cites; AS2 does not wait for it, and the accounts stay other people's (AS2).
Result
Done on aifoundry1 on 28 Sep at 20:51, and by the owner on aifoundry2 on 30 Sep at 14:37:54: both now offer only public-key logins (an ssh attempt without credentials at 14:40:27 got Permission denied (publickey) from both). aifoundry3 has no sshd (U23). The read-only sudo review has not run; the accounts stay with the lab (AS2). Roll back: remove the file and reload ssh.
Command
printf 'PasswordAuthentication no\nKbdInteractiveAuthentication no\n' \
  > /etc/ssh/sshd_config.d/10-labfix-nopassword.conf
sshd -t && systemctl try-reload-or-restart ssh.service
sshd -T | grep -Ei '^(passwordauthentication|kbdinteractiveauthentication) '   # both "no"
# the review, separate and read-only (it can run at any time):
getent group sudo adm; grep -rh NOPASSWD /etc/sudoers /etc/sudoers.d/
Risk to others
Anyone who logs in over the LAN with a password, rather than through Tailscale or with a key, loses that way in: tell the users first. The review changes nothing. Roll back: remove the file and reload ssh.
Problems
H19
U23 OpenSSH on aifoundry3, key-only, for when Tailscale fails#

2 · BlocksWaits for the owner's decisionFrom AS3 (a fallback way into aifoundry3)Where aifoundry3

What
aifoundry3 has no sshd, so when tailscaled or the tailnet fails nobody can log in, root included (H24). Install OpenSSH with password logins off before it first starts. On the tailnet address Tailscale SSH keeps answering port 22, so this adds a way in over the LAN only.
Command
mkdir -p /etc/ssh/sshd_config.d
printf 'PasswordAuthentication no\nKbdInteractiveAuthentication no\nPermitRootLogin no\n' \
  > /etc/ssh/sshd_config.d/10-labfix-nopassword.conf
apt-get install -y openssh-server
sshd -T | grep -Ei '^(passwordauthentication|permitrootlogin) '   # no, no
Risk to others
It opens port 22 on aifoundry3's LAN address, for key logins only: one more service to keep patched, and a way in without Tailscale's check (H1) for anyone whose key is in their ~/.ssh/authorized_keys. aifoundry3 has its own admin, so tell them first. Roll back: apt-get purge openssh-server.
Problems
H24
U24 Peer names in /etc/hosts on every lab machine#

4 · Wastes timeDone on all threeFrom AS4Where all three

What
On all three machines the router's DNS answers for the other two machines' bare names, so an unpinned ssh aifoundryN goes to the Wi-Fi LAN address (H23): from aifoundry1 to a host with no sshd, or to password OpenSSH. Pin the other two machines to their Tailscale addresses.
Result
Done on all three: aifoundry1 on 28 Sep, aifoundry2 and aifoundry3 on 30 Sep; aifoundry3's wrong line (it named itself instead of aifoundry2) was corrected at 15:36 that day.
Command
mkdir -p /root/labfix-20260928 && cp -a /etc/hosts /root/labfix-20260928/hosts.before-peers
for h in aifoundry1 aifoundry2 aifoundry3; do
  [ "$h" = "$(hostname)" ] || printf '%s %s\n' "$(tailscale ip -4 $h)" "$h"
done >> /etc/hosts
getent hosts aifoundry1 aifoundry2 aifoundry3     # the other two: their 100.x addresses
Risk to others
It changes where every user's ssh aifoundryN goes: through Tailscale, which may ask for its browser check (H1), and no longer to the LAN's password OpenSSH. Anything else that uses the bare names moves to the tailnet too. Tell the users first, and aifoundry3's admin. Roll back: copy the saved file back.
Problems
H23

Site and hardware

U25 Test card 0's link at Gen3 before the visit#

1 · EndangersRan 30 Sep; took aifoundry1 downFrom SH1 (its optional first step)Where aifoundry1 card 0

What
aifoundry1 card 0's link logs about one corrected error a second, idle or not (C15). Retrain it at 8 GT/s for ten minutes with the card idle: if the errors stop, the Gen4 signal is marginal; if they go on, a lane or the slot is at fault. Either way the visit for SH1 knows what to bring. Then set it back to 16 GT/s.
Command
P=0000:00:01.0              # card 0's root port; card 0 is 0000:01:00.0
et-who --check && flock -n /run/lock/etsoc-shire0.lock sh -c "   # card 0 free; hold its lock for the test
  grep RxErr /sys/bus/pci/devices/$P/aer_dev_correctable
  setpci -s $P CAP_EXP+30.w=0003:000f && setpci -s $P CAP_EXP+10.w=0020:0020   # target 8 GT/s, retrain
  lspci -vv -s 0000:01:00.0 | grep LnkSta                                       # Speed 8GT/s, Width x8
  sleep 600; grep RxErr /sys/bus/pci/devices/$P/aer_dev_correctable
  setpci -s $P CAP_EXP+30.w=0004:000f && setpci -s $P CAP_EXP+10.w=0020:0020   # back to 16 GT/s
"
Risk to others
It took the whole of aifoundry1 down on 30 Sep, not only card 0 as this note said before: every session, both cards and the CI runner, until someone power-cycles the host on site (C28).
Result
Run by the owner on 30 Sep from 14:40:21, with card 0 idle and its lock held. The two readings at 16 GT/s, a minute apart, gave 76 corrected receiver errors at the root port and none at the card. The retrain to 8 GT/s at 14:41:21 took the whole host down (C28), and it needs a power cycle on site (SH8). Do not run it again. Its question, whether the errors come from the signal at Gen4 or from a lane or the slot, goes to the visit (SH1).
Problems
C15

Where the 47 requests of 27 September went#

32 stay as they were, 9 stay with a part moved to ours, and 6 moved to ours (28 September)
Request, 27 SepNow
CF2 A clock pin on aifoundry3 that leaves the governor aliveStays. The guard's reset gap it named would be closed by our et-reset (U16), which would remove the marker and restart the guard if aifoundry3's admin agrees; the guard and its pin stay with that admin.
CF3 Recover a hung card without power-cycling the host, and find why aifoundry2's hungStays, narrowed to the firmware and driver questions (why the Master Minion hung; why the sysfs reset did not recover it). Resetting aifoundry2's card is done (U15, 08:32), and the et-reset wrapper is ours (U16); who may run it stays the lab's rule (PO2, CF10).
MO1 Free aifoundry1's diskStays: the space is other people's data. Our part is done: our snapshots were destroyed at 08:33 (U3).
MO2 aifoundry1's corrupt files, and a backupStays: the damaged files are other people's, and a backup of their data is the lab's decision. The scrub that follows is ours (U17).
MO4 Run the health check dailyMoved to ours: U18, the daily health check.
MO5 Reboot aifoundry1 and aifoundry2 into 7.0.0-34Moved to ours: U19, the reboots.
MO6 Quiet the desktop services on the headless hostsMoved to ours: U20, the desktop services; done on aifoundry1 on 28 Sep at 20:51.
MO7 After the campaignMoved to ours: U21, tmux and the timers.
AS2 Passwords, sshd and sudoStays for the accounts, sudo and the adm group, which are other people's. Turning off password logins, and the sudo review, are ours (U22); password logins are off on aifoundry1 since 28 Sep 20:51.
AS3 Tailscale: check mode, and a way in when it failsStays for Tailscale's check mode and the invite (the tailnet admin). OpenSSH on aifoundry3 is ours (U23).
AS4 Peer names on every lab machineMoved to ours: U24, the peer names in /etc/hosts; done on aifoundry1 on 28 Sep at 20:51.
SH1 aifoundry1 card 0: cooling first, then a reseatStays for the cooling and the reseat (on site). Its optional first step, the Gen3 test, is ours (U25).
DI5 The PCIe DMA engineStays for the DMA engine's documentation. The IOMMU-passthrough boot is ours (U26).
PO3 A standing maintenance windowStays: the window binds every user. Running the upgrades, reboots and checks in it is ours, as root.
PO5 Headless machinesMoved to ours: U27, the headless target; set on aifoundry1 on 28 Sep at 20:51, from its next boot.
CF1, CF4, CF5, CF6, CF7, CF8, CF9, CF10, CF11, CF12, RT1, RT2, RT3, RT4, RT5, RT6, MO3, AS1, SH2, SH3, SH4, SH5, SH6, SH7, DI1, DI2, DI3, DI4, DI6, PO1, PO2, PO4Stay as they were (AS1 now adds that closing the root path need not end the owner's root).

The six that moved keep their old links: they now lead to our items.

Where the 25 September list went#

The 22 items of the old section 3, “What the Nekko team can do, in order”, and what became of each
25 Sep itemNow
Card lock for everything automatedkept as PO2 (with H31); the CI and demo part is MO3, and the demo's network exposure is in AS1
Free aifoundry1's diskkept as MO1; our part, the snapshots, is done (U3, destroyed on 28 Sep at 08:33)
Close the easy security gapskept as AS1 and AS2 (other people's accounts and sudo); turning off password logins is ours since 28 Sep (U22)
Finish the rebootsours since 28 Sep (U19, after the headless target, U27); aifoundry3 was done on 25 Sep
Write the onboarding pagewe drafted it (U12); adopting it is DI1
A one-page sheet per cardwe drafted it (U12); adopting it is DI2
Recovery cheaper than a power cyclekept as CF3, now the firmware and driver questions (why aifoundry2's card hung, C27, and why the sysfs reset did not recover it); the reset itself is done (U15, the management reset at 08:32 on 28 Sep) and the et-reset wrapper is ours (U16); the shared supply is in SH2
Two small fixes: aifoundry2's /etc/hostsours since 28 Sep (U24), for all three machines; done on aifoundry1 on 28 Sep at 20:51
Two small fixes: the Gen3 test of card 0's linkours since 28 Sep (U25), the step before SH1's visit: a no-reboot Gen3 retrain with card 0 idle. Because the errors continue with the card idle, the test can tell a marginal Gen4 signal from a lane or slot fault
Package the ET driverkept as CF4
One reference /opt/etkept as RT2
Let every tool address one cardkept as CF7
Make the management path robustkept as CF5
Controlled root helperssplit: et-reset is ours (U16), the clock pin stays in CF3, the privilege split is CF10
Firmwaresplit: thermal protection CF1 (new), provenance CF9, signing CF12, telemetry and traps CF11
Toolchain, simulator and librariessplit: RT3 (libraries), RT4 (toolchain), RT6 (simulator)
A firmware policykept as PO1, now with the two upstream fixes
A standing maintenance windowkept as PO3
Power and the sitesplit: SH1 (card 0), SH2 (power), SH3 (console visit), SH4 (Ethernet), SH6 (IP-KVM)
Storagekept as PO4
An access policy, written downmerged into AS1 (root paths) and AS3 (check period); the fallback into aifoundry3 when Tailscale is down is ours (U23)
A lab health checkwritten by us as a read-only command (et-lab-health, U8) and installed on all three hosts on 28 Sep; running it daily is ours (U18). The per-card firmware query was left out because it opens the card
Measurement hygienesplit: logging the room is in SH5; pausing the timers is ours (U21); recording et-lab-manifest and pausing our own cron are ours (section 4), so they were dropped from this list

New since 25 September: CF1, CF2, CF6, CF8, RT5, SH5, SH7, DI4, DI5, DI6, the wider scope of CF5 and CF10, and on 28 Sep CF3's firmware and driver questions about aifoundry2's hung card. Two more, the daily health check and the desktop services, are ours since 28 Sep (U18, U20).

3. Deep dives#

Each problem opens with its state on 27 September: its status, since when, the evidence, what differs per host, and who acts next. Below that is the 25 September write-up, kept as history (new problems have only the current one): what a user sees; what went wrong and why, with each cause marked known, inferred or unknown; how we found it; what it cost; what was done or the workaround; and what the Nekko team can do. Evidence links go to the repository; names in this style are working files of the investigation (section 5).

3.1 Cards, driver and firmware#

What the four ET-SoC-1 cards, their kernel driver, their firmware and the vendor's management tools do that a user would not expect. 33 problems on 4 October: 8 fixed, 2 partly fixed, 13 open, 10 new; 8 of the open and partly fixed ones updated since 27 September.

C1 aifoundry1's two cards were refused by every ET tool: the driver module had an empty version string#

Fixed 25 Sep blocks work aifoundry1

4 October Fixed 25 Sep Since 25 Sep 15:02 PDT (us, root)

Holds through the 30 Sep and 2 Oct boots: all three hosts load the driver with its version string (0.20.0, the same srcversion), and aifoundry1's card nodes came up at the 2 Oct boot. What is left is the lab's: package the driver so it cannot recur, and a clearer error upstream.

Next: Nekko team CF4, CF5; us U18

The 25 September write-up, kept as history:

Who acts (25 Sep): Done. Lab: package the driver. Nekko: clearer error

What a user sees

Every ET tool (dev_mngt_service, et-powertop, and any program built on the runtime) died right after opening /dev/et0_mgmt or /dev/et1_mgmt with Error unable to evaluate compatibility!; dev_mngt_service aborted with terminate called … FATAL SIGNAL RECEIVED and left crash reports. tools/etcfg, which uses a raw ioctl, still answered, so the cards looked half alive.

What went wrong, and why

known The loaded et_soc1 module had an empty version: /sys/module/et_soc1/version held only a newline, where aifoundry2 and aifoundry3 read 0.20.0. DKMS on aifoundry1 built it from /usr/src/et-soc1-0.20.0, a stale snapshot of et-platform 09531e5c1 (21 Oct 2025) whose Makefile has the typo $(ET_MODULE_VERSION=). Upstream fixed it five days later in 78ed9b0d6 ("Fix version on dkms build"). deviceLayer's openWhenReady (DevicePcie.cpp L270–294, linked statically into every tool) reads that file, finds no x.y.z, and throws before any command reaches the card. DKMS had already rebuilt the same broken module for the next kernel, so a reboot would not have fixed it.

Our first diagnosis (22 Sep: "srcversion mismatch, libDM.so refuses the card") was wrong: nothing checks srcversion, and the check is not in libDM.so. unknown What loaded the old module 3 min 42 s after boot.

How we found it

22 Sep: every tool failed and we recorded it as a srcversion mismatch (E21). 25 Sep 13:01–13:15: a read-only investigation for Roman's question "what is broken on 1?" (sysfs, modinfo, hashing the DKMS source against et-platform history); fixed as root the same afternoon.

What it cost

Two of the lab's four cards were unusable from at least 18 Sep (the first crash reports) until 25 Sep 15:02, and inferred plausibly since the broken source was installed in October 2025. Every measurement of ours before 25 Sep ran on two cards, and a wrong diagnosis stood in our notes for three days. A user without root could not have fixed it.

What was done

25 Sep 15:01–15:02 PDT, as root, with the CI runner paused and nobody on a card: dkms remove et-soc1/0.20.0 --all; the old source and the three installed .ko files moved to /root/et-soc1-fix-20260925/; /usr/src/et-soc1-0.20.0 refilled from git -C /usr/src/et-platform archive 353f20e et-driver (the fixed source was already on the machine); dkms install for 7.0.0-30, -31 and -34; modprobe -r et_soc1 && modprobe et_soc1. The module now reads 0.20.0 / 47D26A305A0428B29FB7FC4, byte-identical to aifoundry2 and aifoundry3, for the running and the next kernel. A firmware-revision query succeeded on both cards as an ordinary user at 15:02; kernels ran on both cards from about 16:40, and our queues have run on both since 17:13 (still 0.20.0 at 17:18).

How to verify

cat /sys/module/et_soc1/version prints 0.20.0. Not yet seen: the first boot of aifoundry1 into 7.0.0-34 (H6).

How to roll back

The commands in section 1 of the troubleshooting report, with the backups in /root/et-soc1-fix-20260925/.

What the Nekko team can do
  • Ship the ET driver to every lab host as one versioned package (a .deb with DKMS, built from a tagged et-platform commit) instead of hand-copied /usr/src snapshots, with a post-install check that fails loudly when the module version is empty.
  • Upstream: make deviceLayer's error say what it read and where ("driver version '' in /sys/module/et_soc1/version is not x.y.z"), and make dev_mngt_service catch the exception instead of aborting.
  • A lab health check (cron or the login banner) that prints the module version and the result of one read-only query per card, so a card that cannot be opened is noticed within hours.

Requests: CF4, CF5

Evidence

Re-check, 27 September

  • aifoundry1 27 Sep 21:29: et_soc1 0.20.0 / 47D26A305A0428B29FB7FC4, the 7.0.0-34 build identical; /dev/et* 0666, not recreated since 25 Sep 15:02; card 1 ran our queues through 27 Sep 20:25

Back to the table · Requests

C2 Nothing reliable loaded the ET driver at boot on aifoundry1 and aifoundry3#

Fixed 30 Sep corrupts results aifoundry1, aifoundry3

30 September Fixed 30 Sep Since 30 Sep 15:07 (tested at boot on all three)

Tested at boot on all three: after the power cycle of 30 Sep at about 15:07 each host loaded et_soc1 0.20.0 by itself and created every card's nodes, aifoundry1 for the first time since its fix of 25 Sep.

  • aifoundry1 configured, untested: /etc/modules-load.d/et_soc1.conf in place, but no reboot since 18 Sep (up 9 days on 7.0.0-31 at 23:20 on 27 Sep): the 7.0.0-34 reboot (ours now, U19) is its first test
  • aifoundry3 fixed: held through the 25 Sep 16:20 reboot: driver at +6.2 s, guard on attempt 1; the guard's marker matches the current boot
  • aifoundry2 not affected: always had the entry

Next: Nekko team CF4, DI4; us U19

The 25 September write-up, kept as history:

Who acts (25 Sep): Done. Tell aifoundry3's admin

What a user sees

On aifoundry3, if the demo web service is disabled or the Wi-Fi is slow at boot, the card comes up without its 600/400 MHz, 0 W pin (C5), and results assumed to be at 600 MHz silently are not. On aifoundry1 something unidentified loaded et_soc1 3 min 42 s after boot.

What went wrong, and why

known Neither host had an /etc/modules-load.d entry, and the module has no PCI alias, so udev does not load it. On aifoundry3 the only loader was etsoc1-demo.service's ExecStartPre, ordered after network-online (Wi-Fi) and tailscaled, while the clock guard polls for /dev/et0_mgmt 30 times at 2 s. On the 18 Sep boot the module loaded at +14.65 s and the guard succeeded on attempt 5 at +15.9 s; the demo had failed once at that boot ("Cannot assign requested address") and recovered 3 s later. unknown What loaded it at +3 min 42 s on aifoundry1: there is no root cron, and the CI workflows do not.

How we found it

The aifoundry1 investigation, and the root audit of aifoundry3 (the order of the guard and the demo in the boot journal), 25 Sep.

What it cost

None realised: a silent failure waiting for a reboot with slow Wi-Fi or a disabled demo.

What was done

aifoundry1, 15:02:41: /etc/modules-load.d/et_soc1.conf. aifoundry3, 16:15:36 (fix B1): the same file, plus the drop-in /etc/systemd/system/et-board-clock-guard.service.d/10-labfix-load-module.conf (Wants=/After=systemd-modules-load.service, ExecStartPre=-/usr/sbin/modprobe et_soc1). The guard's own logic is unchanged. aifoundry2 already had the entry. After aifoundry3's reboot into 7.0.0-34, et_soc1 loaded at +6.2 s, the guard set the card on attempt 1 (16:21:06), and its marker holds the new boot_id followed by 600 400 0.

How to verify

cat /run/et-board-clock-guard.ok /proc/sys/kernel/random/boot_id on aifoundry3: the same boot_id, then 600 400 0. On aifoundry1, lsmod | grep et_soc1 right after its next boot.

How to roll back

rm /etc/modules-load.d/et_soc1.conf (and on aifoundry3 the drop-in), then systemctl daemon-reload.

What the Nekko team can do
  • Ship the modules-load.d file with the driver package (C1), and order any service that configures a card after systemd-modules-load.service.
  • Tell aifoundry3's admin about the drop-in (not done yet).

Requests: CF4, DI4

Evidence
  • labfix/audit-aifoundry3.md F3; labfix/consistency.md A5
  • labfix/hostlogs/aifoundry3/W3-config.log (B1 at 16:15:36)
  • troubleshooting report, section 6
  • read-only check 25 Sep 17:17: aifoundry3 runs 7.0.0-34, et_soc1: loading at 6.2 s, marker = boot_id

Re-check, 27 September

  • aifoundry1 and aifoundry3: modules-load.d entry present; aifoundry3 guard drop-in 10-labfix-load-module.conf present

Back to the table · Requests

C3 Stale and foreign ET driver builds sat in DKMS, a risk at every kernel update#

Fixed 25 Sep wastes time latent all three

27 September Fixed 25 Sep Since 25 Sep 16:14-16:25 (us, A2/A3)

Holds. One install recipe (the package) is still the lab's.

Next: Nekko team CF4; us U18

The 25 September write-up, kept as history:

Who acts (25 Sep): Done. Lab: one install recipe

What a user sees

aifoundry1's dkms status listed an old esperanto driver (September 2025 source, the same PCI IDs) for 23 kernels, 20 of them not installed, with WARNING! Diff between built and installed module!. aifoundry3 carried a hand-swapped et-soc1.ko for a removed kernel, signed with aifoundry2's key. Headers of kernels that were not installed made DKMS build drivers for kernels that do not exist, and each host had 33–49 leftover (rc) kernel packages and up to 36 orphan /lib/modules directories. On aifoundry1 the fixed driver had been built by hand in two source trees but never registered with DKMS, so each kernel update left it behind.

What went wrong, and why

known Drivers were installed ad hoc by several people over time: DKMS packages from different snapshots, hand-built and hand-copied .ko files, and old kernels never purged. The risk: if the stale source stopped compiling against a new kernel, DKMS would fail that kernel's postinst and leave the update half-installed, which is what happened on 24 Sep for another reason (H5).

How we found it

The dkms status in the aifoundry1 fix transcript (15:00), then the root audits of all three machines.

What it cost

None realised; a latent cause of half-installed kernel updates.

What was done

Fixes A2 and A3 (aifoundry3 16:14–16:15, aifoundry2 16:23, aifoundry1 16:24–16:25): dkms remove esperanto/0.20.0 --all with its source and 50-esperanto.rules; aifoundry3's stale 6.17 module; 22 stale build products in aifoundry2's DKMS source tree; the headers of kernels that are not installed; kernel 7.0.0-30 on aifoundry1; the rc packages; the orphan /lib/modules directories. At 17:19 dkms status on all three lists only et-soc1/0.20.0 for 7.0.0-31 and 7.0.0-34, and /lib/modules holds only those two kernels. (A transient Could not locate dkms.conf for esperanto on aifoundry1 at 16:26 had cleared.) Each host keeps two rc entries, systemd-timesyncd and apport-core-dump-handler: the packages that chrony and systemd-coredump replaced at 16:14–16:26, whose configuration stays for a rollback.

How to verify

dkms status shows only et-soc1 for the installed kernels; modinfo -k 7.0.0-34-generic -F srcversion et_soc1 gives 47D26A305A0428B29FB7FC4.

How to roll back

Backups in /root/labfix-20260925/ (usr-src-esperanto-0.20.0.tgz, var-lib-dkms-esperanto.tgz, lib-modules-orphans.tgz, lib-modules-6.17.0-40-updates-dkms/, et-soc1-src-build-products.tgz): untar, then dkms add and dkms install.

What the Nekko team can do
  • One documented way to install the ET driver (the package in C1), never a hand-built or hand-copied .ko.
  • After every kernel install, check dkms status and modinfo -k <new kernel> -F version et_soc1 in the health check.

Requests: CF4

Evidence
  • labfix/audit-aifoundry1.md F17, F18, R2, R7; labfix/audit-aifoundry3.md F15; labfix/consistency.md A2, A3
  • labfix/hostlogs/aifoundry1/A-session-20260925T162431.log (A3 section)
  • read-only checks 25 Sep 17:18–17:19 (dkms status on all three)

Re-check, 27 September

  • all three 27 Sep: dkms status = et-soc1/0.20.0 for 7.0.0-31 and -34 only; /lib/modules = -31, -34; no kernel installed since

Back to the table · Requests

C4 Four cards run three firmware releases and idle at different operating points; the driver reports the same nameplate for all#

Open updated corrupts results aifoundry1 card 0 (1.4.1), card 1 (1.2.0); aifoundry2, aifoundry3 (1.3.1)

4 October Open updated Since 22 Sep (mix); 27 Sep (governor differences found in the source)

No sign of a reflash: /opt/et and the driver are unchanged (the firmware query was not run, since it opens the card). The dates of the /dev/et* nodes, which earlier checks used, are no evidence either way: they change at snap refreshes (aifoundry3 at 07:17 and aifoundry2 at 22:43 on 2 Oct, with nothing but snap and AppArmor in the kernel log). The kernel log is the check.

27 Sep: Wider: the three releases also differ in governor behaviour and known bugs, which strengthens the case for one release with both upstream fixes.

27 Sep: the three releases now also differ in governor behaviour and bugs: card 1 (1.2.0) appears to run with DVFS off (C24); card 0's 0.21.x governor does nothing while the card is idle and its 0.21.0 base has a uint16 power overflow above 65.535 W (fixed upstream 26 Nov 2024; unknown in 0.21.2); 1.3.1's safe state may not move the PLL (C22). 1.4.1's 'low_power' is a label for board power ≤ 30 W, not a governor state. This strengthens the case for one release with both upstream fixes.

30 Sep: from the firmware source (read for E54, untested), the releases also differ in how they set a rail's voltage: 1.2.0 writes each NoC set to flash, and 1.3.1 loops until the watchdog resets the card after a failed set; the upstream fix, 7c6049087, is in none of the lab's builds (C12, PO1).

Next: Nekko team PO1, DI2, DI3

The 25 September write-up, kept as history:

Who acts (25 Sep): Lab (Roman): firmware policy

What a user sees

The same program behaves differently per card. aifoundry1's card 0 idles in a low_power state at 300 MHz, 398 mV and 18.8 W; its first 2 s launch after that idle drew no extra power and reached 600 MHz only at the end (t = 4.3 s), after which it idled at 600 MHz and 26 W. Card 1 idles at 600 MHz, 499 mV and 33–35 W. aifoundry2 and aifoundry3 idle at 600 MHz (32 W on aifoundry2 at 74 °C). The driver's configuration ioctl reports TDP 65 W and a 600 MHz boot clock for every card, so it shows none of this. A rule such as "drop any sample off 600 MHz" drops every burst on card 0.

What went wrong, and why

known Different flashed releases (firmware component versions):

aifoundry1 card 0aifoundry1 card 1aifoundry2 and 3
Release1.4.11.2.01.3.1
BL1 / BL20.21.20.18.00.20.0
PMIC1.6.11.3.01.5.0
Minions0.24.00.22.00.23.0
Idle statelow_power, 300 MHz, 398 mV, 18.8 W600 MHz, 499 mV, 33–35 W600 MHz (aifoundry2 32 W at 74 °C; aifoundry3 pinned, C5)

unknown Why the lab carries three releases. Card 0 also reported a peak-hold die temperature of 123 °C before we had run anything on it, and its error counters show ThermThrottleCeEvent 1 and PmicCeEvent 2 (17:18); since it overheats under load (C21), the 123 °C is inferred probably real, not a glitch.

How we found it

The first management query after the aifoundry1 fix (15:02), then a telemetry config read and a two-launch clock test on each card (about 16:40).

What it cost

A 75 KB amendments file of per-card rules before aifoundry1's data could be used; without an idle-aware rule, card 0 came out INSUFFICIENT on all 16 latency items in synthetic passes. A new user comparing cards sees idle power differ by 30–40% and a slow first launch, with no explanation.

Where it stands

Worked around by us for our own data: amendments written before any aifoundry1 data record each card's idle state and judge only samples inside kernel windows, and every result names the card and its firmware. Since 16:26 the aifoundry1 login banner lists each card's firmware (fix A12). Unifying the firmware is a lab decision (fix plan C3.3): flashing with esperanto_flash_tool can brick a card, and it makes that card's earlier data non-comparable.

What the Nekko team can do
  • Choose one supported release for the lab (1.3.1 is the majority) and reflash in a maintenance window with a recovery path, or keep the mix and publish a per-card sheet: release, PMIC, TDP policy, temperature threshold, idle state, boot clock, known quirks.
  • Release notes that describe 1.4.1's low_power idle state and what changed between 1.2.0, 1.3.1 and 1.4.1.
  • Make the management CLI print the firmware's effective TDP next to the nameplate value.
  • 27 Sep: Pick a release that contains e024210bc (safe state) and 478275330 (power overflow), and state each card's active-power-management setting.

Requests: PO1, DI2, DI3

Evidence

Re-check, 27 September

  • No reflash seen on any host: /opt/et unchanged since 25 Sep, /dev/et* not recreated, no card reset since 25 Sep 16:22 (aifoundry3)
  • card 1 read firmware 1.2.0 in our tel block 26 Sep 04:16; banners still list 1.4.1 / 1.2.0 / 1.3.1 / 1.3.1
  • 27 Sep source reading (heatplace/firmware.md): card 1 (1.2.0) appears to run with DVFS off (C24); card 0's 0.21.x governor does nothing while idle and 0.21.0 has a uint16 power overflow above 65.535 W (fixed upstream 478275330, 26 Nov 2024; unknown in 0.21.2); 1.3.1's safe state may not move the PLL (fixed upstream e024210bc)

Back to the table · Requests

C5 aifoundry3 is held at 600 MHz by a root boot service, not by its firmware, and the host cannot see it#

Open updated workaround corrupts results aifoundry3

27 September Open updated workaround Since 25 Sep (pin); 27 Sep (latch found)

Same pin, new consequence: aifoundry3 has no working thermal governor (C23).

27 Sep: the same pin also silences the governor. With TDP 0 the SP's power-down loop cannot exit at the 600 MHz bottom point, and after the first 65 °C crossing the card latches: no governor line and no thermal step until the SP reboots (inference from the source; every trace since 25 Sep is empty, E41 found no throttle/idle event where ≥ 5 were predicted, and the heat-placement experiment's trace probe found it silent at 16:52 PDT on 27 Sep). aifoundry3's slower SP pass (224 ms vs 133 ms on aifoundry2, same firmware) may be that spinning task. The guard is still in place (the probe saw 600 MHz, resting at 53 °C).

Next: Nekko team CF2, DI2

The 25 September write-up, kept as history:

Who acts (25 Sep): aifoundry3's admin (the policy); Nekko (the ioctl)

What a user sees

aifoundry3 never runs above 600 MHz (all 7,745 samples of one session), idles about 11 W below aifoundry2, reads 0.92–0.95× aifoundry2's power in every table, and logs Power throttle down event, current pwr 35380 tdp level: 0 at every kernel start. The driver's configuration ioctl reports TDP 65 W, as on the other cards. Our notes called it "a flashed zero".

What went wrong, and why

known Root's et-board-clock-guard.service, installed on 23 Jul 2026 by aifoundry3's admin, runs /usr/local/sbin/configure-et-board-clock at every boot: DM_CMD_SET_MODULE_STATIC_TDP_LEVEL -l 0, then DM_CMD_SET_FREQUENCY 600,400. Its stated reason: this card "becomes unreliable when firmware DVFS raises the minion clock above its 600 MHz minimum". With TDP 0 the governor's step-down test (power > TDP) is always true and its step-up test never is. The TDP is a RAM value in the service processor (g_pmic_power_reg.module_tdp_level), while the driver's GET_DEVICE_CONFIGURATION returns the nameplate 65 W. The guard writes /run/et-board-clock-guard.ok (<boot_id> 600 400 0). unknown Whether the card really is unreliable above 600 MHz: that is the guard author's claim, untested by us.

How we found it

22 Sep (E21) we read TDP 0 from the firmware and assumed it was flashed; on 25 Sep, during the aifoundry1 investigation, we found the root-owned unit and its marker.

What it cost

Three days of a wrong explanation in our notes and published pages. Any user who reads the driver's 65 W, or assumes the single-card machines are alike, gets aifoundry3's throughput and power wrong.

The workaround

The pin is deliberate and was left in place. We treat aifoundry3 as pinned (tools/claims-v3/lib.sh) and compare cards on switching power over idle, not absolute watts. Since 16:15 the aifoundry3 login banner states the pin (A12). Checked at 17:17, after the reboot into 7.0.0-34 and a per-card reset: the marker matches the boot_id, and our own config read shows TDP 0 W, 600 MHz and max_power. Still open: our pages still say "flashed" (D18), and the guard does not re-check the card after a reset (C6).

What the Nekko team can do
  • Record every deliberate per-card operating-point policy in one visible place: the banner (done) and the lab page.
  • Upstream: have the configuration ioctl return the firmware's live TDP, threshold and boot clock, or document that it returns nameplate values.
  • If this card really is unreliable above 600 MHz, it may be marginal silicon or cooling; worth a vendor look.
  • 27 Sep: Pin the clock with a cap that leaves the governor alive (a maximum-frequency or VMIN-table limit), not a 0 W TDP: at TDP 0 the governor latches and the card has no software thermal response (C23).

Requests: CF2, DI2

Evidence
  • aif1/facts/aifoundry3/clock-guard-scripts.txt (the guard's text and its stated reason)
  • docs/reports/data/2026-09-25-aifoundry1/facts.md (Side findings)
  • docs/findings/16-dvfs-and-leakage.md (26 throttle-down events, 7,745 samples at 600 MHz)
  • et-platform 836a4ab: device-bootloaders/src/ServiceProcessorBL2/services/thermal_pwr_mgmt.c l.576–590
  • read-only check 17:17: marker = boot_id e43d7668…; our tel-smoke config read: tdp_w 0, 600 MHz

Re-check, 27 September

  • aifoundry3 27 Sep 21:19: /run/et-board-clock-guard.ok = boot_id e43d7668... 600 400 0; heat block p3101 telemetry minion 600, NoC 400 MHz
  • NEW: with TDP 0 the SP power task cannot leave POWER_DOWN at the 600 MHz floor; after the first 65 C crossing the governor latches (E41: no throttle/idle event where >= 5 predicted; heat R0 probe 16:52 PDT 27 Sep SILENT)

Back to the table · Requests

C6 After a card reset, aifoundry3's clock guard trusts its boot marker and does not re-check the card#

Open corrupts results latent aifoundry3

30 September Open Since 23 Jul (script); seen 25 Sep

Unchanged: the guard script is as it was on 23 Jul, and aifoundry3's card has not been reset since 25 Sep 16:22. et-reset is installed on aifoundry3 since 30 Sep 14:39 (U16); it never touches the guard and prints a reminder to tell aifoundry3's admin instead, so after a reset the 600 MHz pin still has to be restored by that admin (CF2, DI4).

Next: Nekko team CF2, CF3, DI4; us U16

The 25 September write-up, kept as history:

Who acts (25 Sep): aifoundry3's admin (the guard's author)

What a user sees

Nothing yet. If a per-card reset ever returns the card to firmware DVFS, every "aifoundry3 is pinned at 600 MHz" assumption breaks with no warning.

What went wrong, and why

known From the guard script: configure-et-board-clock exits with verified from boot marker whenever /run/et-board-clock-guard.ok matches the boot_id and the target point, without querying the card (to avoid draining a trace backlog). A per-card reset does not change the boot_id. On 25 Sep our reset test (16:22:21) was followed by a guard restart (16:22:29) that logged verified from boot marker … tdp=0W without reading the card. The card was in fact still at 0 W and 600 MHz at 17:14–17:17, per our queue's reads, so the pin held this time; the guard would not have noticed if it had not. unknown Whether a reset always preserves the setting (one test).

How we found it

Reading the guard script after the reset test's log line (25 Sep).

What it cost

None so far; a latent source of silently unpinned data.

Where it stands

Nobody has changed the guard.

What the Nekko team can do
  • Make the reset procedure remove /run/et-board-clock-guard.ok before restarting the guard; or key the marker to something a reset changes (the device's reset count, the service processor's uptime); or read the TDP once, which is cheap.
  • Document: after any card reset, rm /run/et-board-clock-guard.ok; systemctl restart et-board-clock-guard.

Requests: CF2, CF3, DI4

Evidence
  • labfix/hostlogs/aifoundry3/W3-post.log (the reset; 16:22:29 verified from boot marker)
  • aifoundry3: /usr/local/sbin/configure-et-board-clock (the marker short-cut before any query)
  • our queue's reads on aifoundry3 at 17:14 and 17:17 (tdp_w 0 after the reset)

Re-check, 27 September

  • /usr/local/sbin/configure-et-board-clock unchanged since 23 Jul (sha256 c3317a749b8eb598); no card reset since 25 Sep 16:22

Back to the table · Requests

C7 On cards with a free governor the clock follows the die temperature#

Open updated workaround corrupts results measured on aifoundry2 only (from a cool die); aifoundry1 card 0 unmeasured, but it throttled 10 times on 25 Sep

30 September Open updated workaround Since 27 Sep (scope narrowed)

Now measured on aifoundry2 many times: in DV2's validation (28–29 Sep, E51) its governor entered the thermal loop above a 65 °C mean and left it at 65 °C with no dead band (51 entries, 52 exits), acted on an idle card too, and climbed from 600 to 800 MHz in one step. Its die idles at 71–76 °C, so outside a cool spell it stays at 600 MHz (H28). Card 1 never steps (C24), aifoundry3 is pinned and latched (C23), and card 0 was not measured; its driver counted 10 thermal-throttle events on 25 Sep (C21).

The 'free governor' cards are fewer than the report says. aifoundry1 card 1 never changed clock in 359,657 samples, including 11,446 below 65 °C and readings up to 88 °C, and its trace probe was silent on 27 Sep 20:24 PDT: its DVFS appears off (C24). aifoundry2's governor works but its die never read below 65 °C during the campaign, so it sat at 600 MHz in all 336,070 samples (H28). Card 0 is not used, so its governor was not measured; its driver's counters show that it does throttle (10 thermal-throttle events on 25 Sep 16:39–17:46, C21). By 27 Sep the governor had been seen following temperature only on aifoundry2, from a cold start (once, E10); DV2's validation has since measured it there many times (above). From the source (27 Sep): the thermal input is the integer mean of 34 sensors, the 1000-tick step delay probably lasts about 0.4 s on the cards (the SP timer runs at 10 MHz, FreeRTOS assumes 4 MHz), and 0.18.0 steps in 50 MHz, not LUT points.

Next: Nekko team CF3, CF8, SH5

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko: a supported way to pin a card

What a user sees

The same kernel gives different watts, TFLOPS or cycles-per-ns from one run to the next. Some telemetry samples read 700 or 800 MHz although the card "runs at 600". Zeros ran at 11.4–11.8 TFLOPS against 9.3 for random data, only because the clock differed; random data at 800 MHz touches 88 W for an instant.

What went wrong, and why

known From the cards' firmware source (cafe03fc3^, the closest match to release 1.3.1, C8) and from telemetry. The governor is thermal first: while a kernel runs it steps the minion clock down when the whole-degree mean of 34 shire sensors reads above 65 °C or SoC power exceeds the 65 W TDP, and up otherwise, once per service-processor pass, with no dead band. Below about 65–68 °C a busy card climbs to 700–800 MHz in the middle of a burst and drops back when the master minion goes idle. Nothing in the launch API reports the operating point, and there is no supported user control for a fixed one; aifoundry3's pin uses a zero TDP as a workaround (C5).

How we found it

The minion clock in every 10 Hz telemetry sample; the seven cool-start runs of E10.

What it cost

The 23 Sep 12:51–13:10 rerun session on aifoundry2 (die at 65 °C) had 700–800 MHz in 5–25% of the samples of most bursts and was discarded. An earlier conclusion that "the clock never moves" had to be withdrawn. Every power session since spends 3–8× the card time on preheating and waiting.

The workaround

By us: blocks preheat the die to at least 76 °C with 2 s random-fp32 launches (heat_to in tools/claims-v3/lib.sh), and the analysis drops any burst whose busy samples read off 600 MHz (an idle-aware version on aifoundry1). The aifoundry2 banner says its card reaches 800 MHz only from a cool start.

What the Nekko team can do
  • Put the governor's rules in the lab notes and the banner: 65 °C on a whole-degree mean, 65 W TDP, and on these cards' build instantaneous SoC power.
  • Give users a supported, reversible way to pin a card for one reservation: for example a root helper, run through sudo while the user holds the card lock, that sets DM_CMD_SET_FREQUENCY and the TDP and restores the firmware policy on release. aifoundry3's guard already does the setting part.
  • Have the runtime report the minion clock at the start and end of each launch; add a temperature dead band; publish each card's VMIN table.

Requests: CF3, CF8, SH5

Evidence

Re-check, 27 September

  • aifoundry1 card 1: 600 MHz in all 359,657 v3 samples incl. 11,446 below 65 C and readings to 88 C; R0 trace probe 20:24 PDT 27 Sep SILENT (C24)
  • aifoundry2: never below 65 C in 336,070 samples, so always 600 MHz (H28); gs queue 26 Sep still had off-600 bursts in p11/p21/p31 (re-run by design)

Back to the table · Requests

C8 The firmware on the cards is a mid-2024 build, but the source everyone reads is from December 2025#

Open updated workaround corrupts results all four cards

30 September Open updated workaround Since 27 Sep (mapping extended by us)

We mapped all four cards to their source, and the mapping, with each build's governor, has been in the repository since 28 Sep (docs/findings/14-card-behaviour.md); 1.4.1's exact source is still not public (CF9).

27 Sep: we mapped the other two releases. 1.2.0 (BL2 0.18.0, card 1) is et-platform da192816a (“Close development of version 0.18.0”, 27 Mar 2024), and the closest public source for 1.4.1 (0.21.2, card 0) is 50310b06b (“Close development of version 0.21.0”, 25 Sep 2024): 0.21.1 and 0.21.2 are not in the public history. The PVT driver is identical in all four versions. The mapping and each version's governor went into the repository on 28 Sep (commits 8eb0e38 and a8b3984).

Next: Nekko team CF9, PO1

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko: publish release-to-commit

What a user sees

Statements derived from the firmware source did not match the cards: aifoundry3's trace lines exist only in firmware older than 60b40c10f (24 Sep 2024), and the clock climbed to the top operating point in one step, which the 353f20e governor does not do.

What went wrong, and why

known Release 1.3.1 reports bootloader 0.20.0 and minion runtime 0.23.0, which corresponds to et-platform development between 29 April and 17 May 2024 (closest source cafe03fc3^). The lab's /usr/src/et-platform, and our reading, were at 353f20e (28 Dec 2025). thermal_pwr_mgmt.c has 9 commits in between (685 lines added, 805 removed): for example the cards' build throttles on instantaneous SoC power where 353f20e uses the PMIC average, and it applies a 1.05× TDP guardband on leaving a step-down. The release number appears nowhere in the source tree, so the mapping has to be rebuilt from component versions. aifoundry1's 1.4.1 (bootloader 0.21.2) and 1.2.0 (0.18.0) have not been mapped at all.

How we found it

Mismatches between the cards' trace buffer and the 353f20e source (22–24 Sep); settled on 25 Sep by querying the firmware revisions and bisecting component versions.

What it cost

Every source-derived claim about the governor had to be re-checked against a second source version, and one published statement ("the loop's power input is the PMIC average") was wrong for these cards.

The workaround

By us, 25 Sep 07:25: the mapping to cafe03fc3^, the governor extracted from that commit, and every governor claim re-checked (firmware.md). Not done for 1.4.1 and 1.2.0.

What the Nekko team can do
  • For each firmware release shipped to the lab, publish the et-platform commit and build it came from; tag releases in et-platform (for example fw-1.3.1), and return the commit in the firmware-revision command.
  • Put the release-to-commit table on the lab page, so users read the right source.
  • 27 Sep: Publish the source (or at least the commit) of 0.21.1 and 0.21.2.

Requests: CF9, PO1

Evidence

Re-check, 27 September

  • 27 Sep: 1.2.0 (BL2 0.18.0) -> et-platform da192816a; 1.4.1 (0.21.2) -> closest public 50310b06b (0.21.1 and 0.21.2 are not in the public history); the PVT driver is identical in all four versions (heatplace/firmware.md §1)

Back to the table · Requests

C9 Users cannot run changed firmware: the images must be signed#

Open blocks work all four cards

27 September Open Since 20 Sep

Unchanged. The firmware changes the measurements need are rungs 11–16 of the hub's ladder.

Next: Nekko team CF12

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko / the vendor

What a user sees

Any firmware-side fix or instrument can be built and tested in sys_emu but not run on a card: a counter-select syscall, exporting the 35 temperature sensors, input-side PMIC power, faster unfiltered power, the ECC error sources.

What went wrong, and why

known The boot chain checks signatures, and the open tree has no signing key.

How we found it

Designing the observability improvements, 20–23 Sep.

What it cost

It blocks the firmware rungs of our observability ladder, including the firmware fixes proposed for C10, D1 and D2.

Where it stands

Our PMC-configure syscall patch (patches/0003-pmc-configure-syscall-353f20e.patch) is verified in sys_emu and waits for a signed image.

What the Nekko team can do
  • Offer a path: the vendor signs community firmware builds on request, or the lab keeps one card with a development key for firmware experiments.

Requests: CF12

Evidence

Re-check, 27 September

  • No change possible from a host; nothing new

Back to the table · Requests

C10 A management tool killed mid-request poisons the card's management queue, and samplers often fail to start#

Open workaround blocks work all hosts (seen on aifoundry2 and aifoundry3)

27 September Open workaround Since 24 Sep

No incident since 25 Sep 17:00, and no vendor fix; the banners keep the never-kill-9 rule and the drain line.

Next: Nekko team CF5, DI1

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko: driver / deviceLayer / libDM

What a user sees

After a telemetry sampler or any management client is killed in the middle of a command (kill -9, or timeout's SIGKILL), the next process to open /dev/et0_mgmt (dev_mngt_service, a sampler, the runtime tools) dies with std::bad_function_call, and so does every retry. Separately, a sampler started right after another management client exits comes up with no output about one start in three.

What went wrong, and why

known mechanism, inferred root cause. The killed process's reply stays in the card's management completion queue. The next opener receives a reply to a request it never sent, dispatches it through an empty callback and crashes, leaving its own reply behind, so each crash re-poisons the queue. Neither the driver nor the library discards replies addressed to a process that has gone, and the driver's error counters do not move. The failed starts: unknown cause; inferred the previous holder's release of the node has not finished when the next process opens it.

How we found it

24 Sep: the first start of experiment E32 failed on both cards after a sampler had been killed mid-request. Empty telemetry files in back-to-back runs, 22–24 Sep.

What it cost

It stopped E32's first start on both cards; an unattended runner that simply retries loses telemetry for the rest of its run. Before our runners learned to retry, about one pass in three of the reruns had no sampler.

The workaround

By us: one /opt/et/bin/dev_mngt_service -m DM_CMD_GET_MODULE_POWER -n 0 -u 5000 drains the queue (it crashes on the stale reply and clears it). Our sampler finishes its current request on SIGTERM; the runners stop samplers with SIGTERM only, wait for the first telemetry line, retry up to six times and drain once after a failed start. Since 25 Sep 16:15–16:26 the login banner on all three hosts says never to kill -9 a card tool and gives the drain line (A12). The behaviour itself is unchanged.

What the Nekko team can do
  • Flush a node's management completion queue on close (or open), or tag replies with the opener and drop, with a warning, any reply the library did not issue.
  • Have dev_mngt_service and et-powertop finish their request on SIGTERM and SIGINT; document that timeout -s KILL around a management tool is unsafe.
  • Make open() wait for the previous holder's release, or return a distinct error that a tool can retry on.

Requests: CF5, DI1

Evidence

Re-check, 27 September

  • no bad_function_call on aifoundry1 since 25 Sep 17:00; aifoundry2 last Mgmt kernel line 23 Sep; banners keep the never-kill-9 rule and the drain line

Back to the table · Requests

C11 Nobody could see who held a card node, which only one process can open#

Fixed 25 Sep wastes time all three

27–28 September Fixed 25 Sep Since 25 Sep 16:15-16:26 (us, A12)

Holds. Its follow-up H30 was fixed on 28 Sep (et-who --check, installed on all three); C26 remains. Upstream: name the holder, and a read-only telemetry path.

Two follow-ups of ours, both new on 27 Sep: et-who's idle sentence misled one of our scripts, and it always exited 0 (H30): fixed on 28 Sep, when et-who --check was installed on all three hosts (U2); and our card-1 heat blocks hold card 0's management node with a read-only temperature guard (C26), because a card's temperature can only be read through its single-opener node (C25). The upstream fix for the second is a read-only telemetry path (CF6).

Next: Nekko team CF5, CF6, PO2; us U13

The 25 September write-up, kept as history:

Who acts (25 Sep): Done. CI and demo owners: take the lock. Nekko: holder PID

What a user sees

A second open of /dev/etN_mgmt or /dev/etN_ops fails with Device or resource busy, and nothing tells the user who holds it. A telemetry sampler blocks every other management user (dev_mngt_service, et-powertop) for its whole session. Nothing on the machines explained the rule, /opt/et/bin was not on PATH, and only aifoundry3 had a card lock file.

What went wrong, and why

known The driver allows one open per node (et-soc1-pcie.c l.1108 for ops, l.1148 for mgmt: Tried to open same device multiple times, -EBUSY) and says why only in the kernel log, which users could not read (H15). fuser and lsof show other users' holders only to root, because other users' /proc/<pid>/fd is private. aifoundry2's journal held 27 refused opens and 18 Device is resetting messages in its current boot.

How we found it

Collisions with our own sampler (a 31-minute sampler blocked every other management user during the audit); the root audits counted the refused opens.

What it cost

Every "busy" cost the person who hit it a guess or a message to the channel, and our long samplers silently locked others out.

What was done

Fix A12 (aifoundry3 16:15, aifoundry2 16:24, aifoundry1 16:26):

  • et-who: any user can list the holders. It is a no-argument sudo wrapper around /usr/local/sbin/et-holders, which uses fuser and ps and never opens a node.
  • A login banner (/etc/motd and /etc/update-motd.d/60-labfix-et-who) with the current holders, the one-opener rule, never kill -9, the drain line, and each host's card facts.
  • /etc/profile.d/et-soc1.sh appends /opt/et/bin to PATH.
  • Lock files /run/lock/etsoc-shire<N>.lock (0666, recreated at boot by tmpfiles) with the convention flock /run/lock/etsoc-shire<N>.lock <cmd>, the path aifoundry3's clock guard and CI already used.
  • et-lab-manifest prints the facts a measurement should record (kernel, CPU, power profile, driver, runtime hashes, clock-guard marker).

Still open: the one-opener design itself; aifoundry3's demo does not take the lock (H13); the CI's benchmark scripts take card 0's lock, but not every CI job is known to (H12).

How to verify

As a user: et-who (at 17:18 on aifoundry2 it listed our sampler and block by PID and lock), ls -l /run/lock/etsoc-*.lock shows -rw-rw-rw-, and a new login shows the banner.

How to roll back
rm /usr/local/sbin/et-holders /usr/local/bin/et-who /usr/local/bin/et-lab-manifest /etc/sudoers.d/60-labfix-et-holders /etc/profile.d/et-soc1.sh /etc/update-motd.d/60-labfix-et-who /etc/tmpfiles.d/labfix-etsoc-lock.conf /etc/motd
What the Nekko team can do
  • Have every automated user (CI workflows, the demo app, anyone's queue) take /run/lock/etsoc-shire<N>.lock.
  • Upstream: name the holder's PID in the EBUSY kernel message or in a read-only sysfs attribute, and consider a management daemon (or a read-only telemetry path) that lets a sampler and control tools coexist.

Requests: CF5, CF6, PO2

Evidence
  • labfix/audit-aifoundry2.md E4 (driver source lines, journal counts)
  • labfix/hostlogs/aifoundry3/A-session-20260925T161401.log l.586–700 (A12 files and VERIFY)
  • docs/findings/14-card-behaviour.md l.72–73
  • check at 17:18 on aifoundry2 as a user: et-who lists our sampler and the lock

Re-check, 27 September

  • all three 27 Sep: et-who works as a user and as nobody; lock files 0666; aifoundry1 journal: 4,984 et-holders calls by our queue since 25 Sep 17:00
  • follow-ups: et-who always exits 0 and its idle sentence misled our starter on aifoundry3 (27 Sep 21:10, H30); our heat guard holds card 0's node without card 0's lock (C26)

Back to the table · Requests

C12 Any user can change a shared card's global state: reset it or change its trace level#

Open updated corrupts results all three

30 September Open updated Since 27 Sep (wider)

Wider than reported: stats reset, thresholds, DVFS-off, TDP and frequency are open to every user too, and, from the firmware source (read for E54 on 28 Sep, untested), so is a rail's voltage: on card 1's 1.2.0 each NoC voltage set rewrites the flash sector the boot voltage comes from, and on the two 1.3.1 cards a failed set loops until the watchdog resets the card (CF10, PO1). Since 28 Sep the banners on all three hosts list the other commands (U5), not yet the voltage set.

More global state reachable by any user: (1) measured 26 Sep (E41 TEL-P7): a telemetry stats reset (ettelem sample --reset-ms, DM_CMD_SET_DM_STATS_RUN_CONTROL) restarts the PMIC rail averages on every card (f(1 s) = 0.93 against 0.45–0.68 predicted), so one user's reset changes every other sampler's rail readings; (2) from the source (27 Sep, untested): SET_MODULE_TEMPERATURE_THRESHOLDS accepts any uint8 with no range check, SET_MODULE_ACTIVE_POWER_MANAGEMENT turns DVFS off, and SET_MODULE_STATIC_TDP_LEVEL / SET_FREQUENCY pin the clock; a user can therefore silently disable a card's only thermal protection. Our heat work (27 Sep) sets the SP log level to WARNING only on a card whose probe is ALIVE and restores it; no card was ALIVE, so no level was changed.

Next: Nekko team CF10

The 25 September write-up, kept as history:

Who acts (25 Sep): Lab and Nekko (command privileges)

What a user sees

A card resets under a running job (Mgmt: Device is resetting, action cannot be completed!), or its service-processor trace configuration changes, because of another user's command.

What went wrong, and why

known for reset and trace level: the device nodes are mode 0666, and state-changing management commands worked from our non-root account. We reset aifoundry2 twice on 18 Sep with DM_CMD_RESET_ETSOC (with the owner's approval, before we had seen the lab admin's request not to), and raised the service processor's log level to DEBUG for a voltage map, restoring it afterwards. inferred Not tested: the frequency and TDP set commands are open to users too, since aifoundry3's root guard uses the same CLI.

How we found it

Our own use; the aifoundry2 audit's journal scan.

What it cost

18 Device is resetting lines in aifoundry2's journal on 18 Sep, 17:00–19:00. The lab admin has asked users not to reset cards, because a software reset can hang one.

Where it stands

No change. The banner does not yet say which commands are global.

What the Nekko team can do
  • Split the management commands by privilege: read-only telemetry for everyone; set, reset and trace configuration for a group or through sudo. Log each state change with the caller's UID.
  • Until then, list the global commands in the banner and ask users to restore what they change.

Requests: CF10

Evidence

Re-check, 27 September

  • /dev/et* still 0666 on all three; banners do not list the global commands
  • E41 (26 Sep): one reader's telemetry stats reset restarts every reader's rail averages (f(1 s) 0.93 on every card)
  • source (27 Sep, untested): SET_MODULE_TEMPERATURE_THRESHOLDS accepts any uint8; SET_MODULE_ACTIVE_POWER_MANAGEMENT turns DVFS off; STATIC_TDP_LEVEL and SET_FREQUENCY pin the clock: any user can disable a card's only thermal response

Root steps, 28 September (verified 08:00–08:02; the steps after that, by their own output)

  • all three, 28 Sep (aifoundry2 07:09, aifoundry3 07:26, aifoundry1 07:54): the installed banner (/etc/motd) equals this host's tools/lab/motd-<host> by sha256 (checked again at 08:00–08:02); the 25 Sep one is kept in /root/labfix-20260928/replaced/etc/motd; it lists the commands that change a card for every user (reset, trace level, the telemetry stats reset, temperature thresholds, TDP, frequency, active power management)

Back to the table · Requests

C13 On the two-card host the stock ET tools open both cards; picking one needs aifoundry1's forked runtime#

Open workaround blocks work aifoundry1

27 September Open workaround Since 22 Sep

New cost, from our own tools: a stock tool for card 1 collides with anything on card 0.

Cost seen in the campaign (26 Sep): because the stock dev_mngt_service opens every card's management node, aifoundry1's telemetry passes (E41) had to take the whole host with a host-level lock, and a queue-drain on one card counts as an intrusion on the other card's idle cycle (the block exits 3 and is retried).

Next: Nekko team CF7, DI4; us U13

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko: upstream card selection (ask account rehan)

What a user sees

On aifoundry1, /opt/et/bin/dev_mngt_service -n 1 still opens card 0's management node (and -n 0 opens card 1's), so it fails with EBUSY while the other card is in use, and it holds the other card while it runs. A dead or not-ready card 0 would break tools aimed at card 1. The banner's drain line (-n <N>) has the same limit. The two cards cannot DMA to each other (peer-to-peer DMA not supported).

What went wrong, and why

known from source: deviceLayer's DevicePcie constructor runs openWhenReady on every /dev/etN_mgmt and _ops node whatever -n says (DevicePcie.cpp L486–503), and the runtime's init sends abortDevice and checkDeviceApi to every card (RuntimeImp.cpp L143–150). aifoundry1's /opt/et is a build of a fork (account rehan, acc7ed25) whose deviceLayer adds ET_DEVICES=<n>: open only card n and present it as device 0. Only programs built against it on aifoundry1 get this; the stock dev_mngt_service, identical on all hosts, does not. The fork also assigns the DRAM size to onPkgDRAMBaseAddress_, which inferred looks like a slip; its effect has not been checked.

How we found it

Planning per-card queues on aifoundry1 after the fix (25 Sep); confirmed by reading deviceLayer and the runtime.

What it cost

Design time for per-card queues and a deferred drain. Without the fork's ET_DEVICES, two processes could not use the two cards at once.

The workaround

By us: our tools are built on aifoundry1 against its deviceLayer and run with ET_DEVICES=<n>, so each card has its own queue, and the management drain waits while the other card's lock is held. Two-card runs have worked (the clock test at about 16:40; both queues since 17:13). The fix plan (C2.3) says to keep aifoundry1's /opt/et for now: replacing it would break this and make aifoundry1's data non-comparable. A new user with the stock CLI still hits the problem.

What the Nekko team can do
  • Upstream ET_DEVICES (or honour -n) in the stock deviceLayer, and rebuild dev_mngt_service and et-powertop with it; ask the fork's author (account rehan) to upstream it and to check the onPkgDRAMBaseAddress_ line.
  • Until then, say in the aifoundry1 banner that the stock CLI touches both cards.

Requests: CF7, DI4

Evidence

Re-check, 27 September

  • aifoundry1 26 Sep 00:00-04:19: 13 "Tried to open same device multiple times" on card 0 and 10 dev_mngt_service SIGABRT cores: our stock dev_mngt_service call for card 1 opened card 0 while our watcher held it (C26)
  • E41 had to take the whole host with a host-level lock on aifoundry1

Back to the table · Requests

C14 A hung card is recovered by power-cycling the whole host, although a per-card reset exists and worked on an idle card#

Open updated blocks work all three (reset tested on aifoundry3)

4 October Open updated Since 25 Sep

No hung card since 28 Sep (C27): the restored card ran DV2's 52 validation launches (28–29 Sep) without one. A checked reset command, et-reset, is installed on all three hosts since 30 Sep (U16; aifoundry2's at 15:36), for root and the sudo group. aifoundry2's card is now off the bus for another reason, its cooling (H28, C32): no reset can help it, and none should be tried before the cooling is fixed. Which reset is the supported recovery is still CF3's question, and who may run it the lab's (PO2).

28 Sep: the first hung card since this report. aifoundry2's Master Minion stopped taking work at 02:50:53 PDT during our DVFS development run (C27); the service processor still answered. We did not reset it overnight (the lab's rule). With the owner's approval, the per-card reset through sysfs, which the 25 September write-up below proposes, ran at 06:39:27: it re-attached the card, but launches at 06:41–06:47 still failed with Couldn't use the HPSQ. The management reset (dev_mngt_service -m DM_CMD_RESET_ETSOC -n 0) at 08:32:45 recovered it, and a test kernel ran 3 launches at 08:33 (U15). So a hung Master Minion can be recovered without power-cycling the host, but by the management reset, not the sysfs one. Our et-reset would wrap it with the lock checks (U16); why the sysfs reset falls short, and which reset is supported, is CF3's question.

Next: Nekko team CF3, PO3; us U16

The 25 September write-up, kept as history:

Who acts (25 Sep): Lab: adopt and document the reset. Nekko: firmware abort

What a user sees

After a TensorSend/TensorRecv or a blocking credit wait (csrw fcc) that never completes, every later launch fails with KernelLaunchCmIfaceMulticastFailed or Couldn't use the HPSQ. Perhaps the Master Minion is hanged?, and the card stays unusable until it is reset. The lab norm (lab admin, 17 Jul) is not to reset a card yourself, because a software reset can hang one, and to ask the admin to power-cycle the machine.

What went wrong, and why

known for the hang: a minion has one peer-to-peer ready bit, not one per partner (core-et dcache_reduce.v, partner_ready_peer), so a minion that receives readies from two TensorSend partners at once hangs for good; a blocked credit wait cannot be interrupted; the firmware's abort recovers neither. sys_emu tracks each partner separately, so a schedule that passes there can hang silicon (D9). For recovery: besides DM_CMD_RESET_ETSOC, the driver has a per-card reset, /sys/bus/pci/devices/<bdf>/soc_reset/reinitiate (root, write-only). Neither is documented for users.

How we found it

The nocbench design review against the RTL; the lab norms from the Discord; the aifoundry1 root audit found the sysfs reset in the driver (et_sysfs_soc_reset.c).

What it cost

Each hang means waiting for the admin and a host power cycle that kills everyone's work on that machine.

Where it stands

Tested once by us as root on aifoundry3's idle card at 16:22: echo 1 > /sys/bus/pci/devices/0000:02:00.0/soc_reset/reinitiate. The driver re-enabled the device at 16:22:21, /dev/et0_* came back with mode 0666, every later run worked, and aifoundry3's pin survived (C6). DM_CMD_RESET_ETSOC also worked twice on aifoundry2 on 18 Sep. Not yet tried on a hung card, nor through sysfs on aifoundry1 or aifoundry2, and not available to users. We avoid the hang itself: rings only, partners changed only across a barrier, and a non-blocking credit poll with a bailout before any blocking wait.

What the Nekko team can do
  • Try the sysfs reset on the next hung card before power-cycling the host. If it works, offer it the way et-who is offered: a fixed sudo wrapper et-reset <N> that refuses while the card is held and re-applies the host's clock policy afterwards. Put it in the banner.
  • Vendor: make the firmware's abort recover a hart stuck in a tensor or credit wait, or document which hangs need a reset; model the one-ready-bit rule in sys_emu.

Requests: CF3, PO3

Evidence
  • labfix/hostlogs/aifoundry3/W3-post.log (the reset: enabling device at 16:22:21, nodes back)
  • docs/getting-started.md l.202–209 (lab norm; wedged-card messages), l.305–311
  • docs/et-soc1-notes.md (Rules for kernels that talk)
  • labfix/audit-aifoundry1.md F7 (soc_reset/reinitiate)

Re-check, 27 September

  • sysfs soc_reset present (root-only) on all three; tried once, on aifoundry3's idle card, 25 Sep 16:22; no hung card anywhere since
  • aifoundry2, 28 Sep 02:50:53–02:51:15 PDT: a kernel that never ran, then Couldn't use the HPSQ. Perhaps the Master Minion is hanged? on both later launches; no reset attempted (C27)

Root steps, 28 September (verified 08:00–08:02; the steps after that, by their own output)

  • aifoundry2, 28 Sep: the sysfs reset at 06:39:27 left the Master Minion hung (launches 06:41–06:47 failed); the management reset at 08:32:45 recovered it (a test ran 3 launches at 08:33) (C27, U15)

Back to the table · Requests

C15 aifoundry1 card 0's PCIe link logs about one corrected error per second#

Partly fixed updated wastes time aifoundry1 (card 0, 0000:01:00.0; root port 0000:00:01.0)

4 October Partly fixed updated Since 18 Sep boot; the flood gone since 30 Sep; a trickle since 2 Oct

The flood has not come back, but the count is not zero. Since aifoundry1's boot at 12:59 on 2 Oct (after the fan visit) card 0's root port has counted 336 corrected receiver errors in 46.7 hours, about 7 an hour, against about 3,600 an hour before 30 Sep; card 0 itself and card 1's port count none, and both links run at 16 GT/s x8. The kernel logs them at 2–17 lines an hour. Our dashboard shows this port's rate as 0 an hour, which is wrong (ours to fix). Watch it; on site only if the rate grows.

2 Oct: About one corrected receiver error a second at card 0's root port for twelve days. Since the cold power cycle of 30 Sep at about 15:07 it has counted none (0 at 07:00 on 2 Oct, after 40 hours): the flood was most likely a bad link training at the 18 Sep boot. Our Gen3 test (U25) hung the host (C28) and must not be repeated.

Next: us watch it (U18), fix the dashboard's rate; on site only if it grows

The 25 September write-up, kept as history:

Who acts (25 Sep): On site: reseat card 0

What a user sees

Root port 0000:00:01.0 counts about one corrected receiver error (RxErr) per second: 709,340 in the current boot at 17:18. Card 1's port and the other hosts' ports show 0–3. Card 0 works, but each error is a link-level replay.

What went wrong, and why

Location known, physical cause inferred. The daily count jumped about 250× at the 18 Sep reboots (from about 160–440 a day to 93,000–101,600 a day), and the arrivals are Poisson-like, not periodic firmware activity. The link trains at Gen4 16 GT/s x8 on an x16-capable port. 1.0–1.4 errors per second means a bit error rate of at least about 1e-11, about ten times the PCIe target: likely marginal seating, power or signal integrity at 16 GT/s. If it escalates to an uncorrectable error, the driver disconnects card 0 until reboot, and because the stock tools open every card (C13), tools aimed at card 1 break too.

How we found it

The aifoundry1 investigation traced the 279 MB kern.log to AER lines from port 00:01.0 (25 Sep).

What it cost

Replays on card 0's link (a small DMA cost, not quantified); the log flood (C16); a risk of losing both aifoundry1 cards to one link failure.

Where it stands

The log flood is stopped (C16); the link is unchanged: RxErr 706,250 at 16:26:44, 709,340 at 17:17:59 and 713,050 at 18:07:49, about 1.0–1.2 per second. Not yet tried: the no-reboot Gen3 test (the bandwidth-control cooling device for port 00:01.0), which needs card 0 idle.

What the Nekko team can do
  • On site: reseat card 0 and its power connector, or move it to another slot (then update the BDF in labfix-aer-ratelimit-et0.service), in the same visit as the check of its cooling (C21). Afterwards track TOTAL_ERR_COR per hour over at least two boots.
  • Quick diagnostic without a reboot: drop the link to 8 GT/s and see whether RxErr stops.
  • Card 0's earlier DMA timings are not comparable after a reseat.

Requests: SH1

Evidence
  • troubleshooting report, section 4
  • labfix/audit-aifoundry1.md F4, R4; aif1/verify-hardware.md (rates, timing, error-rate estimate, Gen3 test)
  • read-only check 17:18: /sys/bus/pci/devices/0000:00:01.0/aer_dev_correctable RxErr 709340

Re-check, 27 September

  • aifoundry1 root port 0000:00:01.0 27 Sep 21:29: RxErr 913,006 (TOTAL_ERR_COR 913,018); 0.87/s over 30 s, 1.08/s average since 25 Sep 18:07; BadDLLP 5 -> 12, BadTLP 0 -> 2; nonfatal/fatal 0; 16 GT/s x8
  • card 0 idle since 25 Sep 17:46, yet the rate is unchanged: the errors are not load-driven

Back to the table · Requests

C16 The PCIe error flood filled aifoundry1's logs on a nearly full disk#

Fixed 25 Sep wastes time aifoundry1

27–28 September Fixed 25 Sep Since 25 Sep 16:24 (us, A8)

Holds. On 28 Sep we cleared the systemd reload warning (daemon-reload: 0 units need one) and compressed the flood-era kern.log.1 and syslog.1 (287 and 299 MB to 11.7 and 13.6 MB, content checked). The size cap stays off on purpose.

Next: Us U18

The 25 September write-up, kept as history:

Who acts (25 Sep): Done. Lab: watch the error rate

What a user sees

kern.log and syslog grew by about 570 MB in 5.6 days (four lines per error, 13.7–16.3 KB every 30 s) on a pool that was 95–99% full; the journal kept only about a day (H9), and real driver messages were buried.

What went wrong, and why

known C15's errors, each printed at the kernel's default AER rate limit (5 s, burst 10).

How we found it

kern.log was 279 MB and growing about 850 B/s during the aifoundry1 investigation.

What it cost

Log space on a full disk, and journal retention: the history of aifoundry1's resets and of the 18 Sep jump was lost.

What was done

Fix A8, 16:24:34: labfix-aer-ratelimit-et0.service (enabled at boot) writes 60000 ms and burst 1 to /sys/bus/pci/devices/0000:00:01.0/aer/correctable_ratelimit_{interval_ms,burst}: at most one report a minute; the error counters keep counting. From now on kern.log size says nothing about the link; read aer_dev_correctable instead.

How to verify

16:26:44: 60000 1, kern.log grew 0 B in 30 s, RxErr still rising. 17:19: kern.log grew 618 B in 20 s (about 13,700 B per 30 s before) while RxErr went from 709,400 to 709,418.

How to roll back

systemctl disable --now labfix-aer-ratelimit-et0.service, remove the unit, then write 5000 and 10 back to the two files.

What the Nekko team can do
  • Keep the rate limit until the link is fixed, and put TOTAL_ERR_COR per hour in the lab health check so a worsening link is noticed.
Evidence
  • labfix/hostlogs/aifoundry1/A-session-20260925T162431.log l.22–33 (A8), l.1095–1096 (VERIFY)
  • fix plan A8; read-only check 17:19

Re-check, 27 September

  • correctable_ratelimit 60000 1; unit enabled; kern.log ~37 KB/h (was ~1.6 MB/h); ~1 AER line a minute since 25 Sep 17:00
  • kern.log.1 (287 MB) and syslog.1 (299 MB) rotated 27 Sep stay uncompressed until the 4 Oct rotation (delaycompress; the 25 Sep plan dropped a maxsize cap on purpose)
  • systemctl shows "unit file changed on disk" (NeedDaemonReload) for the unit after the 25 Sep upgrade: cosmetic, clears at the reboot

Root steps, 28 September (verified 08:00–08:02; the steps after that, by their own output)

  • aifoundry1 07:55:30: systemctl daemon-reload (restarts nothing): 294 units flagged before, our AER unit among them; 0 after, no unit changed state, the AER rate limit still 60000 ms, burst 1
  • 07:56:20: gzip -k of kern.log.1 and syslog.1; each .gz passes gzip -t, and its decompressed sha256 equals the original's (6f20f219…, 86e0696e…); owner syslog:adm, mode 0640, mtime kept; the originals then removed
  • logrotate -d rc 0 with the real config; the next rotation renames kern.log.1.gz to .2.gz. /etc/logrotate.d/rsyslog is unchanged (rsyslog's own file): a maxsize 200M line is staged, unused, because a local edit would stall rsyslog's unattended updates and the flood is rate-limited since 25 Sep

Back to the table · Requests

C17 Host programs on aifoundry3 crash about once in 100 launches, 1.08 s after start#

Partly fixed workaround wastes time aifoundry3

27 September Partly fixed workaround Since 26 Sep 04:46-06:15 (our programs rebuilt on aifoundry2 and aifoundry3)

Fixed in our programs on aifoundry2 and aifoundry3 (26 Sep: 641 processes, no crash); aifoundry1's were rebuilt on 28 Sep (six at 07:52, build/sparsity at 20:56), all but the campaign's enercat_v2; the runtime is unchanged, and another account now runs programs on aifoundry3 (H31).

  • aifoundry1 fixed in our programs, 28 Sep: enercat, memhier, memprobe, nocbench, onchip and sgemm rebuilt with the fix at 07:52 and build/sparsity at 20:56 (its host now imports g3log's addLogLevel; the frozen copy is kept as build/sparsity.frozen-hp-20260922); enercat_v2, the campaign's catalogue host, stays as it is (U14); its forked runtime has not shown the crash
  • aifoundry2 fixed in our programs: nine host programs rebuilt ~04:46 26 Sep; runtime unchanged
  • aifoundry3 partly: rebuilt 26 Sep 06:15 except enercat_v2 (the frozen v3 campaign binary, 24 Sep); no crash since 26 Sep 04:11:43

The fix held on the card: aifoundry3's gather/scatter queue (E48), the first card program built with the log-level line, ran 641 host processes (3,096 launches) on 26 Sep 04:48–06:13 PDT with no crash, where the old rate predicts 6.4 (the chance of none is 0.2%). After this report the version-3 campaign's old binaries still lost work to the race on aifoundry3: one memory pass (re-run), five relay processes and one launch. aifoundry2's nine host programs were rebuilt with the fix at about 04:46 on 26 Sep and aifoundry3's after its queue; aifoundry1's follow after the heat work (U13). The runtime itself is unchanged, and the race is in every build.

Next: Nekko team RT1, DI1; us U13

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko (a two-line fix in the runtime); program authors meanwhile (one line at the top of main)

What a user sees

About one launch in 100 of sparsity_host, onchip_host, nocbench_host or enercat_host dies with SIGSEGV (rc 139) 1,078–1,080 ms after it starts, during device setup, on aifoundry3 only. A rerun passes.

What went wrong, and why

Root cause known from four core dumps (25 Sep, 20:22-22:58: enercat_host, onchip_host twice, nocbench_host). The runtime's thread-pool worker (ThreadPool::workerFunc, common-sw/src/threadPool/src/ThreadPool.cpp:81 in et-platform) logs with TP_VLOG(MID), i.e. LOG(logging::VLOG_MID). The custom VLOG_* levels are registered in g3log only by the constructor of logging::LoggerDefault (common-sw/src/logging/include/hostUtils/logging/Logger.h:57-59), which host programs built on the runtime API do not construct. So the first LOG(VLOG_MID) from several freshly started worker threads inserts the level into g3log's global std::map at the same time (g3::logLevel -> std::_Rb_tree::_M_emplace_hint_unique -> _Rb_tree_insert_and_rebalance, the frame of every crash), from different threads, while rt::IRuntime::create is still starting the CommandSender threads: a data race that corrupts the map. The same corrupted map explains the 17:47 abort at exit (glibc heap corruption while g3log destroyed that map). It is timing-dependent: aifoundry3's -O3 runtime build shows it about once in 100 launches; aifoundry2's unoptimized build has not shown it.

How we found it

Our version-3 runs on 25 Sep (rc 139 in 2 of about 20 smoke launches), then the same crash site in every segfault line in the kernel log during the root audit.

What it cost

About 1% of aifoundry3 launches repeated. A new user would read it as a bug in their own code.

The workaround

Our runners count rc 139 as a failed launch and repeat it. Core dumps are now kept on all hosts (H16); four of them (20:22–22:58) gave the root cause above, and the crashes went on after the 16:21 reboot (the latest at 23:04 cost a whole measurement pass). We then reproduced the race without a card: tools/g3log-race in the public repository (commit 27dce7f) lets 4 threads make the first lookup of an unregistered level together, and 6–7% of 20,000 trials corrupt the map (1,087, 1,358 and 1,412 in three runs), against 0 of 60,000 with the level registered first. Every host program in our repository now registers the levels at the top of main; the version-3 campaign keeps the binaries it registered, so it can still crash. The 17:47 abort at exit (glibc heap corruption while libg3log destroyed the same map) fits the same race. The runtime swap (fix plan B5) is not needed.

What the Nekko team can do
  • Fix it in the runtime: register the VLOG_* levels before any thread starts (in rt::IRuntime::create, or a static initializer in the logging library), or make g3log's level lookup non-inserting.
  • Until then, document that a host program must register them before creating the runtime: construct logging::LoggerDefault, or call g3::only_change_at_initialization::addLogLevel(level, false) for VLOG_HIGH, VLOG_MID and VLOG_LOW at the top of main (as our registerRuntimeLogLevels() does).
  • Build the reference runtime with ThreadSanitizer once: it finds this class of race in one launch.

Requests: RT1, DI1

Evidence
  • labfix/audit-aifoundry3.md F1 (six segfault lines, disassembly at +0xd4452) and F2
  • validate3/lessons.md (1 in 100, 1,078–1,080 ms, never on aifoundry2)
  • tools/g3log-race/README.md in the public repository (the reproduction, commit 27dce7f)
  • read-only checks 17:19 (libetrt.so sha256 d8e64390ec428a9b…) and 18:07 (coredumpctl info of our own process: the abort's stack)

Re-check, 27 September

  • E48 on aifoundry3: 641 processes / 3,096 launches with the fix, 0 crashes (6.4 expected; chance of none 0.2%)
  • aifoundry3 after the 25 Sep reboot: 9 SIGSEGV + 3 SIGABRT, all ours, last 26 Sep 04:11:43; none at 23:04 (nearest 22:58:10); the v3 campaign lost mem p3 and five LAT processes to it
  • libetrt.so unchanged on every host (the race is in the runtime)

Back to the table · Requests

C18 Each host runs a different build of the vendor runtime in /opt/et#

Open corrupts results all three

27 September Open Since 22 Sep

Unchanged: the three runtime builds still differ; E50 now shows that host-side timings follow each build.

E50 (27 Sep) shows the runtime builds' effect directly: DMA-only rates agree across hosts, but staged copies follow each host's memcpy and small-copy and launch latencies follow each build's polling constants (D20, H29).

Next: Nekko team RT2

The 25 September write-up, kept as history:

Who acts (25 Sep): Lab (Roman), with account rehan and aifoundry3's admin

What a user sees

Host programs that link /opt/et/lib/libetrt.so run different runtime code on each host: launch and copy overheads and assert behaviour differ, and ET_DEVICES works only on aifoundry1.

What went wrong, and why

known

  • aifoundry2: the stock 2026-01-03 build of 353f20e with an empty CMAKE_BUILD_TYPE (no optimisation).
  • aifoundry3: libetrt.so and libetrt_static.a replaced on 17–23 Jul by a Release -O3 build of 836a4ab plus the event-id fix (about 3.4× fewer instructions, asserts off). Its original libetrt.so is lost: the backup is 0 bytes.
  • aifoundry1: libetrt, libdeviceLayer, the firmware images and 51 test binaries rebuilt on 1–10 May 2026 from a fork (acc7ed25 = upstream f0da105d3 plus two local commits), -O0; plus February-2026 ET copies in /usr/local, one libg3log of them in the linker cache.

libDM.so, dev_mngt_service and et-powertop are byte-identical everywhere. The three RISC-V toolchains are separate builds but generate identical code (D11).

How we found it

The aifoundry1 investigation found aifoundry3's patched library; the consistency check hashed every /opt/et file on all three hosts.

What it cost

Host-side timings are not comparable across hosts.

Where it stands

Partly: aifoundry3's backups were moved out of /opt/et/lib to /var/backups/aifoundry3-opt-et-20260723 (fix B6, 16:15; rollback: move them back); each host's runtime provenance is in its banner; et-lab-manifest prints the libetrt, libdeviceLayer and dev_mngt_service hashes. Not unified: don't replace aifoundry1's /opt/et blindly (C13), and B5 for aifoundry3 turned out not to be needed for its crashes (C17).

What the Nekko team can do
  • Pick one reference /opt/et (a tagged commit, one CMAKE_BUILD_TYPE such as RelWithDebInfo), package it, and install the same bytes on every host. Keep experimental builds in /opt/et-<name> or a home directory.
  • Record the libetrt sha256 with every measurement (et-lab-manifest).

Requests: RT2

Evidence
  • labfix/consistency.md bottom line 2, A7–A16; labfix/optet-diff.txt
  • labfix/audit-aifoundry1.md F9; labfix/audit-aifoundry3.md F2; labfix/audit-aifoundry2.md E5, E7
  • labfix/hostlogs/aifoundry3/W3-config.log (B6: left 0, moved 5)

Re-check, 27 September

  • libetrt.so sha256 prefixes unchanged: aifoundry1 d2412f178ba5434c, aifoundry2 f5d1bfb14038b103, aifoundry3 d8e64390ec428a9b; nothing in /opt/et newer than 25 Sep
  • E50 (27 Sep): DMA-only rates agree across hosts, but staged copies and small-copy/launch latencies follow each host's runtime build and memcpy (D20, H29)

Back to the table · Requests

C19 The vendor tools mislead when something is wrong#

Open updated wastes time all three

4 October Open updated Since 26 Sep 00:02 (aifoundry1 crash file); 27 Sep (source)

Two more misleading readouts found (the maximum temperature reads 0; “low_power” is only a power threshold). apport's coredump hook, which wrote duplicate crash reports of our programs, is off on all three hosts since 30 Sep (aifoundry1 since 28 Sep; U20). On aifoundry2 the hook's failed unit cleared with the reboot, and the host reads running (4 Oct); only apport's own crash report of 28 Sep is left in /var/crash, for root to remove (U20).

  • aifoundry1 regressed on 26 Sep, cleared again on 27 Sep; again on 28 Sep, cleared, and the hook off: the 25 Sep fix (A13) had emptied /var/crash of ET reports; our aborts of 26 Sep 00:02 left _opt_et_bin_dev_mngt_service.1009.crash because apport's hook is still enabled next to systemd-coredump (ours now, U20). We moved it to our home directory at 22:15 on 27 Sep (U4, sha256 unchanged); /var/crash is empty again. On 28 Sep we found that these copies come from apport's systemd-coredump hook (apport-coredump-hook.conf), not from the three apport units the plan named, which are inactive here; switching the hook off is ours now, waiting for the owner's decision (U20), so we left it. /var/crash is still empty. On 28 Sep a crash report of ours from wireplumber (10:27) was there again: we removed it at 21:20 and masked wireplumber in our user manager, and at 20:51 the hook was switched off, with the owner's approval (U20). /var/crash was empty at 01:10 on 29 Sep.
  • aifoundry2 open: the hook is off since 30 Sep 14:37 (U20); its failed unit cleared with the reboot of 15:07, but apport's own report of 28 Sep is still in /var/crash (removing it needs root)

Found in the source on 27 Sep: DM_CMD_GET_MAX_TEMPERATURE returns a value that is never assigned (0), and on 0.20.0 (1.3.1) the power-state query compares milliwatts with the bare number 30, so it never reports low power; on 0.21+ “low_power” just means board power ≤ 30 W. dev_mngt_service still aborts with a core dump on a busy node instead of reporting EBUSY and the holder (10 cores on aifoundry1 on 26 Sep, C26), and the udev rule still matches nothing.

Next: Nekko team CF5, CF1; us U20

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko (et-platform)

What a user sees
  • deviceLayer's Error unable to evaluate compatibility! does not say what it read (C1).
  • dev_mngt_service aborts with an uncaught exception and leaves apport crash files.
  • DM_CMD_GET_FIRMWARE_BOOT_STATUS answers Received incorrect rsp status: -16006 on a working card (aifoundry3), so it cannot serve as a health test.
  • Each dev_mngt_service run drops dev0_traces.txt and dev0_sp_traces_<date with colons>.bin into the current directory; eight of them sit in the root of our own checkout.
  • The udev rule 50-et.rules (KERNEL=="et-soc1*") matches no device node; the 0666 mode comes from the driver itself, and DKMS rewrites the rule at every build.
What went wrong, and why

known From the source and from observation.

How we found it

Along the way: the aifoundry1 diagnosis, the fix plan's verify steps, and the audits.

What it cost

Part of the three-day misdiagnosis of aifoundry1; small clean-ups for every user.

Where it stands

Worked around by knowing it: our health checks use DM_CMD_GET_MODULE_FIRMWARE_REVISIONS or DM_CMD_GET_ASIC_CHIP_REVISION; the stale ET crash reports were moved out of /var/crash (A13); we run dev_mngt_service from a scratch directory.

What the Nekko team can do
  • Print the parsed value in version errors; catch exceptions in the CLI; fix or document GET_FIRMWARE_BOOT_STATUS.
  • Write trace files only when asked, or to a named directory, without : in the names.
  • Fix the udev rule to match et[0-9]*_mgmt and _ops.

Requests: CF5, CF1

Evidence

Re-check, 27 September

  • dev_mngt_service still aborts (SIGABRT, core) on a busy node instead of EBUSY and the holder (10 cores 26 Sep on aifoundry1)
  • udev rule 50-et.rules (KERNEL=="et-soc1*") still matches nothing
  • source: DM_CMD_GET_MAX_TEMPERATURE returns a value never assigned (0); on 1.3.1 get_power_state compares mW with 30, so it never reports low power; on 0.21+ "low_power" only means board power <= 30 W

Back to the table · Requests

C20 The driver's signing key is not enrolled: turning Secure Boot on would make the cards disappear#

Open blocks work latent all three

27 September Open Since 24 Sep

New detail: the firmware is in Setup Mode, so a BIOS reset that restores the default keys could turn Secure Boot on and hide the cards.

Next: Nekko team SH3

The 25 September write-up, kept as history:

Who acts (25 Sep): On site (console)

What a user sees

Nothing today. If a BIOS reset or update turns Secure Boot on, et_soc1 will not load and every card disappears from its host.

What went wrong, and why

known DKMS signs the module with a per-host key (/var/lib/shim-signed/mok/MOK.der) that was never enrolled; Secure Boot is off (Setup Mode), and the kernel logs module verification failed: signature and/or required key missing - tainting kernel.

How we found it

The root audits, 25 Sep.

What it cost

None yet; a lost afternoon on site, with a confusing cause, if it ever happens.

Where it stands

No change: enrolling a key needs someone at the console at the next boot.

What the Nekko team can do
  • Before any BIOS work, at each console: mokutil --import /var/lib/shim-signed/mok/MOK.der, then confirm Enroll MOK at the next boot.

Requests: SH3

Evidence

Re-check, 27 September

  • mokutil --sb-state: SecureBoot disabled; aifoundry1 and aifoundry3 also "Platform is in Setup Mode"; module unsigned by an enrolled key (taint OE)

Back to the table · Requests

C21 aifoundry1's card 0 overheats under load: 98–102 °C in short test runs, 115–117 °C just after#

Fixed 2 Oct updated blocks work aifoundry1 card 0 (0000:01:00.0)

4 October Fixed 2 Oct updated Since 2 Oct 12:59 (fan replaced on site)

Fixed on 2 October: on site, card 0's fan was found broken and replaced (host up at 12:59). Our acceptance test that afternoon: 49 °C idle (the peak since the restart 53 °C) against 65 °C before, and 8 minutes of sgemm bursts (8 s under the lock, 2.5 s gaps) held it at 52–53 °C (peak 56 °C), cooler than card 1 under the same test (59–60 °C, peak 63 °C). The owner put it back in service, and the new-user brief now offers it (SH1). 4 Oct: it has idled at a 48–55 °C mean for two days (peak 57 °C), at 19–21 W, with no error events. The “123 °C hot spot” quoted before was a peak held since September (C30). The installed login banner still warns against card 0 until the corrected one is installed as root (U28).

Firmware side found 27 Sep (from the source, not tested): card 0's 0.21.x governor acts only while a kernel runs, so an idle die at 115 °C gets no software response, and its 0.21.0 base steps the clock up when board power exceeds 65.535 W (uint16 overflow; card 0 drew 66–71 W); whether 0.21.2 carries the upstream fix is unknown. Which path dropped it to 300 MHz (PMIC alarm or idle point) is not established. Card 0 idled at 62–63 °C at 300 MHz during the PCIe runs (27 Sep 14:24–15:13) and has taken no work since 25 Sep; our card-1 heat blocks now watch it (start only ≤ 85 °C, stop > 90 °C). Nothing physical has been checked.

Next: us install the corrected banner (U28, root)

The 25 September write-up, kept as history:

Who acts (25 Sep): On site: check card 0's fan, heat sink and airflow

What a user sees

About ten minutes of short test launches on 25 Sep (our smoke blocks, about 17:45) took card 0's die to 98–102 °C. Right after the last one it read 115–117 °C with nothing running and drew 66–71 W at 600 MHz, until its firmware dropped it to 300 MHz; it then cooled from 104 °C to 78 °C in eight minutes. Card 1, in the same machine, peaked at about 71 °C under the same smokes, and aifoundry2's card stays near 90 °C through hours of work. Idle, card 0 looks normal, because firmware 1.4.1 parks it at 300 MHz and about 19 W between jobs (C4).

What went wrong, and why

Effect known (our telemetry); cause inferred: a cooling fault on this card (its fan, the heat sink or its contact, or the airflow at its slot), hidden by its low-power idle. The 66–71 W with nothing running is leakage at that temperature (D3). Before we had run anything on it, card 0 already reported a peak-hold die temperature of 123 °C, and by 17:18 one ThermThrottleCeEvent. unknown Whether this is related to the same card's PCIe errors (C15); one visit can check both.

How we found it

Our first smoke blocks on aifoundry1's cards after the fix (25 Sep, about 17:45), from the sampler's die temperature and power.

What it cost

Card 0 is excluded from our measurement campaign (amendment A4), so the lab has three cards for sustained work, not four. A user who runs a long job on card 0 gets throttled clocks and inflated power, and inferred risks the card.

Where it stands

By us: card 0 excluded from our campaign (amendment A4, 17:56); a warning in aifoundry1's login banner since 17:56:40 (/etc/motd; the previous text is in /root/labfix-20260925/motd.before-card0-warning; rollback: copy it back). Card 0's smoke data are kept as a record of the fault. Nothing physical has been checked.

What the Nekko team can do
  • On site: check that card 0's fan spins at speed, the heat sink's seating and thermal interface, and the airflow around its slot, with card 1 in the same chassis as the reference; combine the visit with C15's reseat.
  • Afterwards: ten minutes of load at 600 MHz with a sampler running; card 0 should stay within a few degrees of card 1. Then remove the banner warning.
  • Firmware: a card that sits at 115 °C should raise an alarm the host can see, not only throttle.

Requests: SH1, CF1, PO1, CF6

Evidence

Re-check, 27 September

  • aifoundry1 card 0 err_stats (world-readable): PmicCeEvent 18 (board power 75.0-75.75 W against a 75 W threshold), ThermThrottleCeEvent 10, all 25 Sep 16:39-17:46; card 1 all 0; none since (card 0 idle)
  • card 0 idled at 62-63 C at 300 MHz during the PCIe runs (27 Sep 14:24-15:13); our card-1 heat blocks watch it (start <= 85 C, stop > 90 C)
  • source: card 0's 0.21.x governor acts only while a kernel runs, and 0.21.0 steps the clock up above 65.535 W (uint16 overflow; card 0 drew 66-71 W); whether 0.21.2 has the fix is unknown
  • the aifoundry1 banner warning (25 Sep 17:56) is in place

Back to the table · Requests

C22 No card has a die-temperature hard trip or any hot-spot protection, and the lab's firmware releases carry known thermal and power bugs#

New workaround blocks work all four cards (firmware 1.2.0, 1.3.1, 1.4.1)

2 October New workaround Since 27 Sep (firmware source read); 28 Sep (a 103 °C pass found in our 26 Sep data)

Found on 27 Sep in the firmware source of all three releases, and seen in our 26 Sep data (a pass at a 90–103 °C mean, up to 86.9 W, with the clock never leaving 600 MHz). 2 October: an idle card heated to 138 °C (peak sensor reading 144 °C) at 134 W, and nothing on the card acted; it dropped off the PCIe bus at 12:02 (H28). At idle a clock cut cannot help: the power is leakage, which grows with temperature (about doubling every 23 °C above a 13.6 W floor). Only a hard trip that lowers the minion voltage or cuts the minion rails, at about 105 °C, would stop a runaway; until the firmware has one (CF1), a host whose card loses its cooling has no protection.

28 Sep, new evidence from our own data: aifoundry2's version-3 catalogue pass 11 (26 Sep 02:26–02:31 PDT) ran at a 34-sensor mean of 90–103 °C (above 90 °C in 2,565 of its 2,567 samples), with the hottest sensor up to 106 °C and board power up to 86.9 W (145 samples at 75 W or more, the longest stretch 3.2 s). The clock read 600 MHz in every sample and the PMIC's system temperature read 0. So no die trip acted, nothing responded to the hot spot, and the PMIC's 75 W alarm, which should take the clock to 300 MHz, did not show. The pass also broke our own 90 °C rule: that runner had no ceiling (see “Where it stands” below).

Next: Nekko team CF1, PO1; us U1, U13

Who acts: Nekko (firmware); Roman (the firmware policy)

What a user sees

Nothing, until a card overheats. The maximum-temperature query (DM_CMD_GET_MAX_TEMPERATURE) returns 0 on every card. A die can pass 100 °C without an effective die trip: aifoundry1's card 0 throttled (10 events) yet still reached 115–117 °C on 25 Sep (C21), and aifoundry3 and card 1 reach 88–90 °C in our runs with nothing stepping in (C23, C24). And in one of our campaign passes (26 Sep) aifoundry2 ran for 4 minutes at a 90–103 °C mean (hottest sensor 106 °C) with nothing tripping: the clock read 600 MHz throughout, although board power passed the PMIC's 75 W alarm level in 145 samples, up to 86.9 W.

What went wrong, and why

inferred From the firmware source at each card's release (0.18.0 = 1.2.0, 0.20.0 = 1.3.1, 0.21.0 ≈ 1.4.1), read end to end on 27 Sep and not tested on a card:

  • the governor's only thermal input is the mean of 34 shire sensors, each truncated to whole degrees, compared with > 65 and no dead band; the hottest sensor is computed only to be reported (it ran about 3 °C above the mean in E5);
  • max_temp is never assigned in any version, so the maximum-temperature query returns the zero it was initialised with;
  • the PVT controllers' over-temperature interrupts are never enabled;
  • the only hardware trips are the PMIC's alarms, at 75 °C on its own "system temperature" register (the source says the PMIC currently reports it as 0) and at 75 W board power, both to 300 MHz;
  • in 0.20.0 (aifoundry2, aifoundry3) the safe state changes the PLL only if the voltage lookup for 300 MHz fails, while it updates the reported clock either way (fixed upstream in e024210bc, 5 Sep 2024, after these builds);
  • in 0.21.x (card 0) the governor acts only while the master minion is busy, so an idle die gets no software response, and 0.21.0 keeps board power in 16 bits, so above 65.535 W the test wraps and steps the clock up instead of down (fixed upstream in 478275330, 26 Nov 2024; whether 0.21.2 has it is unknown). Card 0 drew 66–71 W at 115–117 °C.
How we found it

Reading the governor code for the heat-placement experiment (27 Sep), before any card time, and checking each path at the release each card runs (C8).

What it cost

No damage is known. It is the likely firmware side of card 0's overheating (C21), and on the two cards whose governor does not act (C23, C24) nothing protects the die short of an unverified PMIC alarm. Our runs stop themselves at a 90 °C mean (one campaign pass on 26 Sep did not: it reached a 103 °C mean on aifoundry2); other users' runs do not.

Where it stands

Worked around by us: our blocks stop at a 90 °C mean (the heat blocks at 80 °C mean or 85 °C on the hottest reading). One runner did not: the version-3 campaign's catalogue runner (25–26 Sep) had no ceiling. Its hot passes heated the die to 88 °C or more before each burst and stopped only for another user or a stopped sampler, which is how aifoundry2 reached a 103 °C mean. aifoundry1's card 0 is never used, and a guard stops card-1 work if card 0 reads above 90 °C. The findings are in our working files and go into the repository after the heat work (U13).

What the Nekko team can do
  • Now, in an hour: say which of the lab's releases carry e024210bc (safe state) and 478275330 (power overflow), which sensor the PMIC's 75 °C alarm reads, and whether it can fire.
  • In firmware: act on the hottest sensor (or the mean plus a margin), enable the PVT over-temperature interrupts as a real die trip, assign max_temp, and report thermal and PMIC events to the host.
  • Reflash to one release that has both fixes (C4's firmware policy).

Requests: CF1, PO1

Evidence
  • heatplace/firmware.md l.20–33 (summary), l.180–246 (the governor loops; the 16-bit power sum at 50310b06b, thermal_pwr_mgmt.c:858), l.250–285 (every clock path; the PMIC alarm, bl2_pmic_controller.h:262,281; max_temp never assigned; the PVT interrupt never enabled, pvt_controller.c:362-372; the safe state at ffca4cbb4, thermal_pwr_mgmt.c:2010-2034)
  • 03-experiments.md E5 (the hottest sensor about 3 °C above the mean)
  • aifoundry2, version-3 catalogue pass 11 (hot condition, hold 88 °C), 26 Sep 02:26:38–02:30:56 PDT: docs/reports/data/2026-09-25-claims-v3/raw/aifoundry2/cat/p11 telemetry.jsonl.gz, reduced with Python on 28 Sep: 2,567 samples at 100 ms; the 34-sensor mean (temp_c.pmic) 90–103 °C, above 90 in 2,565; the hottest sensor (temp_c.minshire[2]) up to 106 °C; board_w up to 86.87 W, 145 samples at 75 W or more, the longest run 33 samples (3.2 s); minion clock 600 MHz and sp.system_c 0 in every sample. The runner's stop conditions: tools/claims-v3/cat/run_catalogue_t10.py l.8

Back to the table · Requests

C23 aifoundry3's 0 W TDP pin also stops its governor, so the card makes no thermal step#

New workaround corrupts results aifoundry3

27 September New workaround Since 25 Sep reset / 27 Sep (found)

Inferred from the source; every service-processor trace on aifoundry3 since 25 Sep is empty, as predicted.

Next: Nekko team CF2, DI2

Who acts: aifoundry3's admin and Nekko: a pin that leaves the governor alive

What a user sees

Every service-processor trace on aifoundry3 since 25 Sep is empty: no throttle, no idle event, no governor line, where a 22 Sep dump had 26 "Power throttle down" lines. Runs heat the die to 88–90 °C and the clock never steps. Its management pass takes 224 ms, against 133 ms on aifoundry2 with the same firmware.

What went wrong, and why

inferred From the source, and consistent with every trace since 25 Sep. With the TDP pinned at 0 W (C5), the power-down loop can exit only when the PMIC's average power is below 1.05 × 0 W or the clock is at 300 MHz. At the 600 MHz bottom point the reduce step does nothing and returns success, so the power task never returns. The first time the mean temperature passes 65 °C, the management task sets a thermal-down request that only the stuck task can clear, so from then until the service processor restarts the card logs no governor line and makes no thermal step. A spinning task would also explain the slower management pass (suggestive, not proof).

How we found it

Designing the heat-placement experiment (27 Sep): the source predicted a latched governor at TDP 0. E41's three telemetry passes (26 Sep) had found no throttle or idle event on aifoundry3 where at least 5 were predicted, and the experiment's trace probe at 16:52 PDT on 27 Sep classed the card silent, resting at 53 °C, as predicted for a latched card.

What it cost

aifoundry3 has no software thermal response, so a hot run keeps heating until something else stops it (ours stop at 90 °C). Any governor or trace experiment on aifoundry3 returns nothing: the heat work could not test its trigger there.

Where it stands

Worked around by us: we treat aifoundry3 as having no governor, cap every run at 90 °C, and compare cards on switching power, not absolute watts. Our banner for aifoundry3, installed on 28 Sep, says that the pinned card makes no thermal step (U5).

What the Nekko team can do
  • Pin the clock with a mechanism that leaves the governor alive (a supported maximum-frequency or VMIN-table cap), not a 0 W TDP.
  • In firmware: let the power task return when the operating point cannot go lower.
  • Confirm once with a service-processor trace at INFO level right after a reboot: the source predicts a stuck power task, then silence after the first 65 °C crossing.

Requests: CF2, DI2

Evidence
  • heatplace/CRITIQUE.md §1.1 (the power task's exit conditions, thermal_pwr_mgmt.c lines 2238–2246, 1788–1791 and 2366–2374)
  • 03-experiments.md E41 (no throttle or idle event where at least 5 were predicted; quiet pass 224.1–224.5 ms on aifoundry3 against 133.2 ms on aifoundry2)
  • the heat-placement probe, 27 Sep 16:52 PDT: class SILENT, rest 53 °C (session log)

Back to the table · Requests

C24 aifoundry1 card 1's clock never moves: its DVFS appears to be off, so it has no thermal step either#

New workaround corrupts results aifoundry1 card 1 (firmware 1.2.0)

4 October New workaround Since 25-26 Sep campaign / 27 Sep (found)

600 MHz in all 359,657 campaign samples, cool or hot, and again in all 10,033 samples of our 29 Sep runs; its trace probe was silent (27 Sep 20:24). 4 Oct: no command reads the active-power-management flag (the management API has only the set, which changes the card for everyone), and 1.2.0's closest public source (da192816a) turns it on at every service-processor boot with no VMIN check, so neither hypothesis below explains a card that never steps. Card 1's service processor restarted at the 2 Oct boot; a busy run on it that afternoon (13:07–13:15, 54–60 °C) logged no clock. One short cool-start run that logs the clock settles whether it steps now (CF8).

Next: Nekko team CF8, PO1, DI2

Who acts: The lab (read the card's power-management flag); Roman (the firmware policy)

What a user sees

Card 1 reads 600 MHz in every telemetry sample: cold, at 45–63 W below 65 °C, and at 88 °C. Its maximum clock since boot is 600 MHz.

What went wrong, and why

inferred In the version-3 campaign (25–26 Sep) it read 600 MHz in all 359,657 samples, including 11,446 below 65 °C, where its 1.2.0 governor should step up in 50 MHz steps to 700 MHz, and readings up to 88 °C, where it should step down. The simplest reading was that active power management is off on this card: set off by SET_MODULE_ACTIVE_POWER_MANAGEMENT, or disabled at boot by an invalid VMIN table. 4 Oct: neither fits the closest public source of 1.2.0 (da192816a), which sets it on at every boot and has no VMIN check (VMIN validation came in April and May 2024); why the card never steps is unknown.

How we found it

Checking which cards could show a governor step, for the heat-placement design (27 Sep). The experiment's trace probe at 20:24 PDT classed card 1 silent (no governor line of any kind), as predicted for DVFS off.

What it cost

The heat work's validation card can show no governor behaviour, and its die reached 88 °C with no step (C22). Our own docs called it governed (D23). The four cards now have four clock behaviours (C4).

Where it stands

Worked around by us: card 1 is treated as fixed at 600 MHz, with runs capped at 90 °C. Our docs are corrected in the working copy (U6), and our aifoundry1 banner says so since 28 Sep (U5).

What the Nekko team can do
  • Read the card's active-power-management flag (15 minutes with card 1 idle; it can be scheduled at any time) and say whether it is off on purpose.
  • If it is not, turn it on, or reflash under C4's firmware policy.
  • Put each card's DVFS state on the per-card sheet (appendix B) and in aifoundry1's banner.

Requests: CF8, PO1, DI2

Evidence
  • heatplace/feasibility.md l.93 (600 MHz in 359,657 of 359,657 samples)
  • heatplace/CRITIQUE.md §1.2 (1.2.0's governor)
  • the heat-placement probe, 27 Sep 20:24 PDT: class SILENT, rest 55 °C (session log)

Back to the table · Requests

C25 The operating system cannot see a card's temperature, or any chassis fan: only the single-opener management node reports it#

New workaround wastes time all three hosts

4 October New workaround Since structural; noticed 27 Sep

Checked on 27 Sep and again on 4 Oct on all three hosts: no ET entry in hwmon, no fan readings. Since 2 Oct the lab dashboard's Live section and the History page show each card's mean temperature once a second, read by our live monitor through the management node (H35).

27 Sep: Checked on 27 Sep on all three hosts: no ET entry in hwmon, no fan readings.

Next: Nekko team CF6

Who acts: Nekko (driver)

What a user sees

sensors and /sys/class/hwmon list the CPU, the NVMe drive, the network adapters and ACPI, but no ET-SoC-1 card and no fan speeds.

What went wrong, and why

known The driver registers no hwmon device, so a card's temperature and power are readable only through its management node, which one process can hold at a time (C11, C13). No Super I/O sensor driver is loaded for the hosts' boards, so the chassis fans are invisible too.

How we found it

The 27 Sep re-check, looking for a way to watch aifoundry1's card 0 without opening it.

What it cost

An overheating card (C21) cannot be watched without taking it from its user; a health check, et-who or the banner cannot warn about temperature; monitors collide with tools (C26).

Where it stands

Worked around: read the temperature inside your own run. The driver's world-readable err_stats counters show board-power and thermal-throttle events after the fact, and et-lab-health, installed on all three hosts on 28 Sep, reports them (U8).

What the Nekko team can do
  • Driver: expose die temperature and board power read-only through hwmon or sysfs, from the service processor's periodic statistics, with many readers allowed.
  • Optionally load the boards' Super I/O sensor driver for fan speeds.

Requests: CF6

Evidence
  • 27 Sep re-check on all three hosts: hwmon names (acpitz, nvme, the NIC, coretemp, Wi-Fi) and no nct67* or it87 module (labreport2/host-aifoundry1.json and the other two)

Back to the table · Requests

C26 Our own tools opened aifoundry1 card 0's management node without card 0's lock#

New wastes time aifoundry1 card 0

4 October New Since 26 Sep 00:00 (watcher); 27 Sep 20:24 (guard, by design)

The heat guards have not run since 28 Sep. But since 2 Oct our live monitor opens every card's management node, card 0's included, about once a second without the card's lock, while et-who --check shows no holder (H35).

30 Sep: 26 Sep: 13 refused opens and 10 aborts. Our heat guard never took card 0's lock: it ran without it during card 1's heat runs of 27–28 Sep and our overheating runs of 28 Sep (10:12–11:58), and the repository's guards still do (U13). No guard has run since.

Next: Us H35, U13; Nekko team CF6

Who acts: Us (after the heat work); Nekko (a read-only temperature path, C25)

What a user sees

On 26 Sep 00:00–04:19 PDT the kernel log on aifoundry1 shows 13 Tried to open same device multiple times for card 0, and dev_mngt_service aborted 10 times with core dumps.

What went wrong, and why

known Our card-0 heat watcher (run from aifoundry2 every 120 s, without card 0's lock) held card 0's management node while our card-1 telemetry blocks ran the stock dev_mngt_service, which opens every card (C13). Since 27 Sep 20:24, the heat blocks on card 1 hold card 0's node with a read-only temperature guard, up to 2,700 s at a time, by design and without card 0's lock, because nothing else can read card 0's temperature (C25).

How we found it

The kernel log and coredumpctl on aifoundry1, in the 27 Sep re-check.

What it cost

Small: card 0 must not take load anyway, the card-1 blocks still ended ok, and the watcher is gone. A tool or CI job that probes card 0 gets EBUSY while the guard runs, and on 26 Sep it broke our own rule that card 0 is never opened.

Where it stands

The guard runs by design during card 1's windows (22:30–23:45 on 27 Sep; 00:00–02:00 and 06:00–08:00 on 28 Sep). After 08:00 on 28 Sep the guard takes card 0's lock for its lifetime, and our drain checks the other card's node, not only its lock (U13).

What the Nekko team can do
  • A read-only, multi-reader temperature path (C25, request CF6) removes the need to open card 0 at all.

Requests: CF6

Evidence
  • aifoundry1 journalctl -k and coredumpctl, 26 Sep 00:00–04:19 (labreport2/host-aifoundry1.json)
  • the heat code's guard, tools/claims-v3/hp/hplib.sh l.247–273 (the heat worktree, not merged yet)

Back to the table · Requests

C27 aifoundry2's Master Minion hung on 28 Sep; the sysfs reset did not recover it, the management reset did#

Fixed 28 Sep blocks work aifoundry2

30 September Fixed 28 Sep Since 28 Sep 08:32 (the management reset, U15)

Recovered on 28 Sep: the per-card sysfs reset at 06:39 re-attached the card but left the Master Minion hung (launches at 06:41–06:47 still failed); the management reset at 08:32:45 recovered it, and a test kernel ran 3 launches at 08:33. It has not hung since: DV2's validation ran 52 launches on it on 28–29 Sep, all returning 0, one of them meeting the same clock step down as the hung launch. The cause, and which reset is the supported recovery, are CF3's questions.

Next: Nekko team CF3

Who acts: Us, with root: the reset, done 28 Sep 08:32 (U15); Nekko: the cause, and which reset is the supported recovery (CF3)

What a user sees

From 02:50:53 PDT on 28 Sep no kernel ran on aifoundry2's card. A kernel launched then never ran (kernel did not finish within 6 s, aborting the stream), and every program after it failed while creating its runtime: Couldn't use the HPSQ. Perhaps the Master Minion is hanged?. The kernel log, 4 s after that launch: ET 0000:02:00.0: Error Event Detected, Level Critical, SP Runtime Error, Runtime Error Count Beyond Threshold: 6. The service processor still answers: at 02:52 it read 600 MHz, 25.9 W at idle and 60 °C, with its threshold at 65 °C.

What went wrong, and why

unknown The cause is not established. What the data show:

  • The hung kernel was the second of four 7 s lifts in our DVFS development run (pass p6041, a heater kernel on all 32 shires). The first lift ran normally at 800 MHz (7.5 s). The second program started 0.6 s after it ended, at 02:50:53.593, just as the governor's idle reset took the clock from 800 to 600 MHz: our sampler's first 600 MHz reading came at 02:50:53.725. The program's 14 ms calibration kernel then ran at 02:50:53.746 and counted cycles at the 800 MHz rate (0.77 GHz by its own count, against 0.58 at 600 MHz in the first lift), so the clock change fell within a few tens of milliseconds of it.
  • The next kernel, the first of the 7 s stream, never ran: board power stayed near the 26 W idle (at most 33.8 W, the tail of the first lift) and the clock at 600 MHz, until timeout stopped the program after 10.1 s.
  • The kernel log's SP runtime error came at about 02:50:57.6 (its boot-relative stamp, calibrated on an audit record's clock), 3.9 s after the calibration kernel ended and before the host gave up at 6 s. It is the sixth such event on this card since the host booted on 18 Sep (counts 1 and 2 on 20 Sep, 3 and 4 on 22 Sep, 5 on 25 Sep). After the first five the card went on running kernels, and the kernel log shows no reset in between, so the event alone does not mean a hang.
  • The next two programs failed after 5.2 s each, inside the runtime's constructor, while it tried to abort the stuck command (RuntimeImp::abortDevice, then abortCommand).
  • inferred Our hypothesis: a launch that meets the governor's idle clock change hangs the Master Minion. Of the night's 46 launches on this card, this was the only one that started at 800 MHz and the only one met by a step down; the 43 launches before it all ran, among them 5 that met a step up (600 to 700 or 800 MHz) within 0.5 s, and 4 that started within 0.7 s of the previous kernel's end. One case is not a finding.
How we found it

Our DVFS development queue on aifoundry2 (28 Sep 00:40–02:54 PDT) stopped itself at the third failed launch; its alert file names the evidence. No threshold, TDP, clock or voltage had been set that night.

What it cost

aifoundry2's card could not run a kernel for almost six hours (02:50–08:32). The owner's DVFS validation, which needs this card, and any other user of it were blocked until it was recovered. The management side still worked (temperature, power and the service processor's trace could be read), so nothing else on the host was affected.

Where it stands

Our queue stopped at 02:54. Following the lab's rule (C14), we did not try a reset overnight; there is no workaround without one. Recovered on 28 Sep at 08:32:45, with the owner's approval (U15), in two steps. At 06:39:27 the per-card sysfs reset (echo 1 > /sys/bus/pci/devices/0000:02:00.0/soc_reset/reinitiate, as root) re-attached the card (the kernel log: enabling device, added peer-to-peer DMA memory) but left the Master Minion hung: our test launches at 06:41, 06:43 and 06:47 still failed (exit status 1; the saved output of the last two shows the runtime's constructor failing with Couldn't use the HPSQ. Perhaps the Master Minion is hanged?). At 08:32:45 the management reset (dev_mngt_service -m DM_CMD_RESET_ETSOC -n 0, sent as our user after et-who showed no holder) recovered it: the kernel log shows Mgmt: Device is resetting, then the card re-enabled and its DMA memory re-added, and the tool reported success at 08:32:52. At 08:33 a 1 s test kernel (our sparsity_host, the fma test on one minion) ran 3 launches, all ok at 600 MHz, and held the card for 1.65 s. The host was not rebooted, so our /tmp files (H22) were never at risk.

What the Nekko team can do
  • Done on 28 Sep at 08:32, by us with the owner's approval (U15): the management reset. The per-card sysfs reset had not recovered the Master Minion.
  • In firmware: find why the Master Minion stopped taking work, starting from the hypothesis above: a kernel launch during the governor's idle reset from 800 to 600 MHz. Say also what the SP's runtime-error count records, and why its five earlier events on this card left it working (CF3).
  • In the driver: say why the sysfs reset re-attaches the card without recovering a hung Master Minion, and which reset is the supported recovery (CF3).

Requests: CF3

Evidence
  • the pass's launch records, the heater's output and telemetry: build/claims-v3/aifoundry2/dv2/p6041/launches.jsonl, heater-1-pre.out.gz (the three errors), tel-1.jsonl.gz (clock and power); the service processor after the hang: p1111/z1.json (02:52); the alert: dv2/ALERT-MM-HANG.json (the DVFS worktree's build directory, not committed)
  • every launch of the night against the clock: labreport2/mm-hang/launch_vs_clock.py and its output (46 launches; 1 started at 800 MHz; 6 met a clock change within 0.5 s, 1 down and 5 up)
  • aifoundry2's kernel log, read with dmesg as the user on 28 Sep at 03:25: labreport2/mm-hang/dmesg-raw-0328.txt; every card error event with its calibrated time: labreport2/mm-hang/kernel_events.py (six SP runtime errors since the 18 Sep boot, the last at 02:50:57.6 on 28 Sep)
  • the development results: dvfs2/DEV-RESULTS.md §1–2 (development data, not validated)
  • the resets and tests of 28 Sep: labreport2/fix28/c27/ (reset-test.out and reset-test2.out: the two failed launches after the 06:39 sysfs reset; mgmt-reset.out: the management reset at 08:32:45; reset-test3.out: the 08:33 test, 3 launches ok); aifoundry2's kernel log (dmesg, as the user) at 06:39 and 08:32; the root session's reset log in /root/labfix-20260928/ on aifoundry2

Root steps, 28 September (verified 08:00–08:02; the steps after that, by their own output)

  • aifoundry2, 06:39:27: the sysfs reset; dmesg: enabling device (0000 -> 0002), added peer-to-peer DMA memory; our test launches at 06:41, 06:43 and 06:47 exited with status 1; the saved output of the last two (reset-test.out, reset-test2.out) reads Couldn't use the HPSQ. Perhaps the Master Minion is hanged? (in RuntimeImp::abortDevice, while creating the runtime)
  • 08:32:45, as our user: dev_mngt_service -m DM_CMD_RESET_ETSOC -n 0: Mgmt: Device is resetting, action cannot be completed! for about 6 s (dmesg -T runs about 10 s ahead of the wall clock), the card re-enabled and its DMA memory re-added; “Service request succeeded” at 08:32:52
  • 08:33:01: sparsity_host, fma, one minion: device ready in 0.17 s, 3 launches ok at 0.599 GHz (534,348 iterations each, 0.487 s), device held 1.65 s

Back to the table · Requests

C28 Retraining aifoundry1 card 0's PCIe link to 8 GT/s took the whole host down on 30 Sep; it needs a power cycle on site#

Fixed 30 Sep blocks work aifoundry1

30 September Fixed 30 Sep Down 14:41–15:07 (our Gen3 test, U25; power-cycled on site)

Our Gen3 test of card 0's link (U25) retrained its root port to 8 GT/s at 14:41 on 30 Sep, and the whole host froze at once: the journal's last entry is at 14:41:17, with no kernel error, machine check or panic record. Roman power-cycled the lab at about 15:07, and it came back with both cards working (the incident and its lesson).

Next: us U25 (never repeat)

Who acts: on site, the power cycle (SH8); Nekko, why a link retrain takes the host down (SH8); us, the checks after the boot, and never repeating the test (U25)

What a user sees

From 14:41 PDT on 30 Sep aifoundry1 answers nothing: Tailscale lists it offline, ssh times out on its tailnet and LAN addresses, and on the LAN it does not even answer ARP (from aifoundry2 at 14:48, ip neigh shows its LAN address as FAILED). Every session on it, both of its cards, its CI runner and any job there went with it, and whatever was in its /tmp is lost at the power cycle (H22).

What went wrong, and why

unknown The cause is not established. What happened: at 14:40:21 the owner, as root, started our Gen3 test of card 0's link (U25), with no holder on either card (et-who --check) and card 0's lock held. Its two readings at 16 GT/s, a minute apart, went through: the root port's corrected receiver-error count rose from 1,146,936 to 1,147,012 (76 in 60 s, the usual rate, C15), and card 0's own count stayed at 0. The next command, at 14:41:21, retrained the root port 0000:00:01.0 with a target of 8 GT/s (setpci on its Link Control 2 and Link Control registers). Nothing came back after it. A read-only command of ours on aifoundry1 had succeeded at 14:41:20; from about 14:41:40 ssh timed out.

  • inferred A host whose Wi-Fi stops answering ARP has hung or panicked; losing card 0 alone would not do that. The kernel on these hosts does not reboot after a panic (kernel.panic is 0 on aifoundry2 and aifoundry3, read on 30 Sep; aifoundry1 is presumably set the same way), so either way it stays down until someone power-cycles it.
  • Section 2.9's risk note for U25 was wrong: it said a failed retrain would drop only card 0, until the next reboot.
How we found it

The owner's terminal stopped after the second reading, and from aifoundry2 we saw aifoundry1 leave Tailscale and the LAN at the same moment.

What it cost

All of aifoundry1, until someone is on site: the sessions logged in there (among them the idle ones of 18 Sep, H27), both cards, the CI runner and any job. The lab has no console or out-of-band access (H21, SH6).

Where it stands

Back since about 15:07, when Roman power-cycled all three machines on site. It booted kernel 7.0.0-34 (H6) to the text target (U27, H8) and loaded the driver at boot (C2); both cards' nodes are there, both links at 16 GT/s x8, card 0's error counters all 0, and its root port counted no corrected errors in the first minutes (C15). No panic record was kept (systemd-pstore found the store empty), which points to a hard hang rather than a panic.

What the Nekko team can do
  • Done on 30 Sep at about 15:07: the power cycle (SH8); section 4.6's checks passed at 15:18–15:20.
  • Read why it went down: done on 1 Oct (SH8): the journal ends at 14:41:17 with no kernel error, machine check or panic record. That boot is now two boots back (journalctl -b -2 -k, root or the adm group), and aifoundry1's journal keeps about ten days (H9), so save it if it is wanted.
  • Decide whether the hosts should reboot by themselves after a panic (kernel.panic), since nobody can reach their consoles (SH6).

Requests: SH8, SH6, SH1

Evidence
  • the owner's terminal: Gen4 14:40:21 RxErr 1146936 RxErr 0 and Gen4 14:41:21 RxErr 1147012 RxErr 0 (the root port's count, then card 0's), then nothing
  • from aifoundry2, 14:41–14:48: tailscale status (offline), tailscale ping (no reply), ping and ip neigh (its LAN address FAILED), port 22 closed on the tailnet and LAN addresses; our audit's last successful command on aifoundry1 at 14:41:20
  • sysctl kernel.panic on aifoundry2 and aifoundry3: 0 (no reboot after a panic), 30 Sep 14:55

Back to the table · Requests

C29 A race in the ET driver: reading a card's message counters while the card is reset, or while the driver loads, can return garbage or crash the reader#

New workaround corrupts results latent all three hosts

4 October New workaround Since driver 0.20.0 or earlier; found 30 Sep, in the source

Found on 30 Sep by reading the driver's source while reviewing our new usage logger; never seen on a card, and not tried. The driver shows these counters before it builds the queue tables they read, and frees the tables before it removes the counters. Our usage logger already ignores an impossible jump. 4 October: filed upstream, publicly, as et-platform issue #136. The chips are end-of-life, so the owner chose a public report. et-platform's head is still 836a4ab of 17 Jul, and the driver on all three hosts is unchanged (0.20.0, the same srcversion). A patch is ready for the maintainers (CF13).

Next: Nekko team CF13

Who acts: Nekko (driver, upstream); until it is fixed, anyone whose tools read these counters

What a user sees

Almost always nothing. Each card's directory in sysfs has world-readable counters of the messages sent to and from the card, mgmt_vq_stats/msg_count and ops_vq_stats/msg_count, with byte counts, rates and a utilisation figure beside them. Tools read them to see whether a card is in use: our usage logger every minute and our lab dashboard every 10 minutes. A read that lands in the brief moment while the driver tears down or rebuilds a card's queues can return stale or absurd counts (a jump of trillions of messages); one that lands while the driver is loading can instead crash the reading program.

What went wrong, and why

known From the driver's source, the same in the installed 0.20.0 and in upstream's latest (et-platform 836a4ab, 17 Jul): the functions behind these files walk the card's queue tables with no lock. When the driver sets the queues up, it records how many there are and shows the files first, and builds the tables after; when it tears them down, it frees the tables first and removes the files after, without clearing the count. Every card reset (the management reset, the per-card sysfs reset, the firmware's MM reset), driver load, unbind and PCIe error teardown passes through one or both of these windows.

  • inferred On a fresh driver load the tables do not exist yet, so a read there is a kernel NULL-pointer fault in the reading program, which would most likely also leave that card's next reset or unload hung until a reboot. During a reset the driver still points at the freed tables, so a read returns whatever that memory now holds: no crash on these hosts' kernels, but garbage counts.
  • inferred Each window lasts well under a hundredth of a second, so a tool reading once a minute has about a 1-in-10,000 chance of a bad reading per reset; one polling in a tight loop during a reset would likely hit it.
How we found it

Reviewing our usage logger before installing it on 30 Sep: the review followed each sysfs file the logger reads into the driver and found the order. We did not try to trigger it: that needs a card reset or a driver reload, and we do not reset cards.

What it cost

Nothing yet: no bad reading, crash or hung reset has been seen.

Where it stands

Open: filed upstream on 4 Oct (et-platform issue #136), and no fixed driver anywhere yet (CF13). Partly worked around in our own tools: the usage logger ignores an impossible jump and starts a new baseline, and it reads the counters only once the card's devnum file exists, which the driver adds only after a load has finished. Neither the logger nor the dashboard yet pauses during a reset.

What the Nekko team can do
  • Driver: show the counter files only after the queue tables exist, and remove them before freeing the tables; the kernel already makes removing a file wait for a read in progress. Also clear the table pointers when they are freed.
  • Ship the fix in the lab's driver package (CF4).

Requests: CF13

Evidence

Back to the table · Requests

C30 The cards report one die temperature: the 35 sensors' own readings stay inside the firmware, and the lab's “hottest sensor” is a peak held since the card started#

New workaround corrupts results all four cards (firmware 1.2.0, 1.3.1, 1.4.1); the lab's dashboard and history page

4 October New workaround Since structural; the dashboard label since 30 Sep, the live readings since 2 Oct; noticed 4 Oct

Found on 4 Oct while looking for a way to break the temperature reading down by sensor. The host gets the integer mean of the 34 minion-shire sensors plus hardware peak-holds kept since the card's service processor started. Our dashboard shows that peak as the “hottest sensor”. In 2–4 Oct's records it fell only once on any card, when aifoundry1 restarted at 12:59 on 2 Oct. Card 1 has shown 63 °C there since 2 Oct while its mean idles at 53–61 °C. Card 0's “123 °C hot spot” at a 65 °C idle, quoted in C21 and SH1, was the peak of its September overheating: the same 123 °C peak at a 64–65 °C mean was read on 25 Sep, on 28 Sep and again on 2 Oct. No command, trace record or debug path gives the sensors one by one.

Next: Nekko team CF14; us the mean’s minimum, average and maximum per history bucket, and the I/O shire’s sensor (the peak was relabelled on 4 Oct)

Who acts: Nekko (firmware); us (the lab's labels)

What a user sees

One number per card. The dashboard reads “57 °C die (hottest sensor 63)”, the live console “hottest 63 °C”, and the history page draws a dashed “hottest” line. ettelem temp prints die_c and die_max_c. ettelem sample prints temp_c.minshire [mean, low, high], temp_c.ioshire [current, low, high], temp_c.pmic (the mean again), sp.minion_c and sp.system_c (always 0). No tool says which shire is hottest, or what any one sensor reads.

The “hottest” number barely moves. On 2–4 Oct the live collectors kept about 35,000 readings per card. Over that time aifoundry1's card 1 showed 63 °C as its hottest from 13:08 on 2 Oct onwards, at every mean from 53 to 61 °C. aifoundry3's card showed 59–61 °C, and aifoundry1's card 0 showed 123 °C until its fan was replaced. The number never fell except at that 12:59 restart.

What went wrong, and why
  • known The die has 35 temperature sensors that the service processor reads: one in each of the 34 minion shires and one in the I/O shire, on five PVT controllers, sampled continuously at 12 bits (PRM §1.6 counts 36, one per tile of the 6 × 6 grid; the open RTL drops the PCIe shire's; PVT 4's mask leaves its other channels off, pvt_controller.c:143-148). Only the service processor can reach them (PRM Table 15-37).
  • known For the host the firmware keeps three numbers (et-platform 353f20e; the PVT driver and this function are the same in 0.18.0, 0.20.0 and 0.21.0). get_module_current_temperature() (thermal_pwr_mgmt.c:732-772) sends: the integer mean of the 34 minion-shire readings, each already truncated to whole degrees (pvt_controller.c:567-585, 1303-1340); the highest and lowest of the 34 sensors' hardware peak-hold registers (:637-640); and the I/O shire's sensor with its own peak-holds. pmic_sys is filled with the mean again.
  • known The peak-holds are cleared only by pvt_hilo_reset() (thermal_pwr_mgmt.c:2828). That runs when the SP's sampling task starts and on a statistics reset (DM_CMD_SET_STATS_RUN_CONTROL, performance.c:573-576), which also restarts the SP's power statistics for every reader. The statistics' minion temperature is a 5-pass running average of the same mean (thermal_pwr_mgmt.c:238, 690-711). Its system temperature is a literal 0 (:703), and the maximum-temperature query returns a value nothing assigns (:1380-1384, C22).
  • known Nothing carries a single sensor. No management command returns one (device_mgmt_api_spec.h:42-140). The driver's per-shire temperature print (pvt_controller.c:1183-1213) and its all-sensor reader pvt_get_and_print() (:1573) have no caller, and the service-processor images shipped in /opt/et do not contain them (we could not check the images flashed on the cards). So the DEBUG trace that carries per-shire voltages (E6) carries no temperatures: aifoundry2's DEBUG captures of 20 Sep hold 34 voltage lines a pass and no temperature line. MDI memory reads are limited to the minions' firmware data and stacks (minion_debug.c:616-623), and the driver exposes no hwmon entry (C25).
  • known Our tools then call the peak the hottest sensor. ettelem temp prints minshire_high as die_max_c (tools/ettelem/ettelem.cpp:189). The live collector records it as max (tools/lab/live/live-collector.py:278, 312). The dashboard prints “hottest sensor” (dashboard/page/script.js:730, live.js:96), and the history page draws the largest peak in each bucket as “hottest” (history/build.py:97, history.template.html:59, 80).
  • inferred Card 0's 123 °C dates from its September overheating. The minshire reading [65, 60, 123] was taken on 25 Sep, and the same mean and peak (65, 123) on 2 Oct, so the lab's 30 Sep power cycle most likely did not restart that card's service processor, although the page calls it a cold power cycle (C15); no reading of card 0 between 30 Sep 15:07 and 2 Oct 10:40 exists. The card's uptime query would settle it.
How we found it

The owner asked on 4 Oct for the temperature readings broken down by sensor. We read the firmware source, the lab's tools and the PRM, and the live collectors' records of 2–4 Oct on all three hosts, all read-only and without opening any card.

What it cost
  • Wrong statements on this page and in the lab's notes: card 0 “65 °C with a 123 °C hot spot” at idle (C21, SH1). Every “hottest sensor” the dashboard has shown since 30 Sep, and the history page since 2 Oct, is a peak since the card started.
  • No measurement can say where on the die heat is, or which shire is hottest. The lab's thermal model stays lumped (the spatial temperature brief).
  • The governor and any trip act on the mean. In 1 s windows the hottest sensor runs 1–4 °C above it (E53). In aifoundry2's runaway the gap grew with the temperature: 1–2 °C below 80 °C, 2–4 °C at 100–119 °C and 6 °C at 138 °C (H28), and nobody can tell which shire that was.
  • The extremes are shared state. One user's statistics reset (ettelem sample --reset-ms) restarts every reader's peaks and the SP's power statistics, and, as our tool sends it, also turns off the SP's statistics trace (performance.c:557-564) (C12).
Where it stands

Partly worked around. The owner's registered experiments read the hottest sensor in each 1 s window by resetting the statistics every second, under the card's lock (E41, E53). Users may not do this: the onboarding brief forbids --reset-ms. What anyone holding the card can read today with ettelem sample, as the onboarding brief does, is the mean, the I/O shire's own sensor, and the highest and lowest any sensor has read since the card started. Done on 4 Oct: the dashboard and the history page no longer call the peak the hottest sensor; they say it is the peak since the card started or its statistics were last reset. Still ours, from the reading the collector already takes: show when the peak last rose; draw the mean's minimum, average and maximum per history bucket; and add the I/O shire's sensor.

What the Nekko team can do
  • Firmware: every sensor's reading, with its place on the die, in a command and in the SP statistics trace; the current hottest sensor and its number in the existing reply.
  • Now, in an hour: confirm that the lab's three releases compute these fields as the open source does, and publish the sensor-to-shire map.

Requests: CF14, CF1

Evidence
  • et-platform 353f20e (unchanged at 836a4ab60), device-bootloaders/src/ServiceProcessorBL2/: driver/pvt_controller.c:143-148, 567-585, 618-665, 1071-1105, 1183-1213, 1274-1340, 1573; services/thermal_pwr_mgmt.c:238, 690-711, 732-772, 1380-1384, 2805-2838; services/performance.c:546-576; rtos_task/dm_task.c:167-276. The same PVT driver and temperature code at da192816a (0.18.0), cafe03fc3^ (0.20.0) and 50310b06b (0.21.0), by git diff: licence header only.
  • PRM §1.6 and Figure 1-3 (p. 18: one temperature sensor per tile), Table 15-37 (p. 400: the PVT controllers are SP-only), the eFuse map (p. 427: calibration for 35 sensors); open RTL core-et/rtl/inc/pvt_defines.vh (35 sensors, the PCIe shire's dropped).
  • aifoundry2, 20 Sep: docs/reports/data/2026-09-20-power-aifoundry2. temp_c.pmic equals the mean in 3,679 of 3,681 samples. The I/O shire's high is held at 90 °C while its current reading falls to 75 °C. sp.system_c is 0. raw/sp1.bin holds 34 “MS n Voltage” lines and no temperature line.
  • The live collectors' records, ~/live/history/2026-10-0[2-4].jsonl on each host, reduced on 4 Oct: aifoundry1 card 0, 35,043 readings: the peak 123 °C at a 59–66 °C mean until 12:59 on 2 Oct, then 53–57 °C at a 48–55 °C mean. Card 1, 35,048 readings: 63 °C before the restart, then 54 °C, and 63 °C again from 13:08:46 that day at a 53–61 °C mean. aifoundry3, 35,146 readings: 59–61 °C at a 52–60 °C mean. One fall in all, at aifoundry1's restart. aifoundry2's runaway, 824 readings (docs/reports/data/2026-10-02-idle-runaway-aifoundry2): the peak 1–2 °C above the mean below 80 °C, 1–3 °C in the 80s, 2–4 °C at 100–119 °C, 6 °C at 138 °C.
  • Card 0's reading of 25 Sep, minshire [65, 60, 123]: our session notes (validate3/lessons.md). The same peak on 28 Sep: 05-claims.md (“standing statistics (never reset)”) and 03-experiments.md. The 1 s-window gaps: E53 in 05-claims.md, and the spatial temperature brief.

Back to the table · Requests

C31 A rejected read-only management query logs “Critical, SP Runtime Error”, although the card is fine#

New wastes time all four cards (seen on aifoundry2)

4 October New Since 20 Sep (the first event); explained 4 Oct, from the source

Found on 4 Oct, reading the service processor's source against the kernel log. Every line the SP logs at error level counts as a runtime error, and with the threshold at 0 each one reaches the host as a Critical event. aifoundry2's first five such events (20, 22 and 25 Sep) each fell in the same second as a batch of our read-only queries, one of which the firmware rejected; the card ran kernels after each. Only the sixth, on 28 Sep, came with a hang (C27).

Next: Nekko team CF3

Who acts: Nekko (firmware)

What a user sees

dmesg shows Error Event Detected, Level Critical, SP Runtime Error, Runtime Error Count Beyond Threshold: N right after an ordinary read-only dev_mngt_service query, and the card goes on running kernels. Nothing says what the error was.

What went wrong, and why
  • known In the service processor's logger (et-platform 836a4ab, ServiceProcessorBL2/services/log.c), every line logged at error level adds one to a runtime-error count. The threshold, RT_ERROR_THRESHOLD, is 0 (mgmt_build_config.h), so every such line sends a Critical event to the host. A command the firmware rejects logs such a line: an unknown command id, or a query the build cannot answer.
  • inferred From the timing: aifoundry2's counts 1 and 2 came at 21:20:18 on 20 Sep, 3 and 4 at 12:10:29 on 22 Sep, and 5 at 07:25:40 on 25 Sep, each in the same second as a batch of our read-only queries. On 22 Sep that batch's residency-throttle query failed with status −9004; the 20 Sep batch had the same query. The card ran kernels normally after each. Only the sixth event, at 02:50:57 on 28 Sep, came with a hang (C27).
How we found it

Matching the calibrated event times in the kernel log of 28 Sep against our own command times, then reading the logger's source (4 Oct).

What it cost

The five events looked like five earlier faults, and they made the 28 Sep hang harder to read (CF3).

Where it stands

Read the line as “the service processor logged an error”, not as a hang, and check whether a launch still runs.

What the Nekko team can do
  • Firmware: put the logged message into the event, or give a rejected command a lower level than Critical.
  • Answer the rest of CF3's question about the 28 Sep hang.

Requests: CF3

Evidence
  • docs/reports/data/2026-09-28-dvfs2-aifoundry2 (incident/kernel-events.out, lines 4–7 and 28)
  • et-platform 836a4ab: device-bootloaders/src/ServiceProcessorBL2/services/log.c lines 274–279 and 414–433; include/config/mgmt_build_config.h line 379

Back to the table · Requests

C32 A card that falls off the PCIe bus still looks present, and a warm reboot does not bring it back#

New workaround wastes time aifoundry2 (1 and 2 Oct); in principle any card

4 October New workaround Since 1 Oct 09:42 (the first drop); noticed 2 Oct

aifoundry2's card dropped off the bus at 09:42 on 1 Oct and at 07:42 and 12:02 on 2 Oct, and has been off since. Each time its /dev nodes stayed, the driver stayed bound, and et-who called the card free; the kernel log showed only queue errors (SQ[0] sync: head mismatched, head_remote: -1), never “card lost”. Only sysfs tells: the card's link speed reads Unknown and its width 63 (still on 4 Oct). A plain reboot at 10:47 on 2 Oct left the slot empty; reboots that cut the slot's power brought the card back.

Next: Nekko team CF5; us H34

Who acts: Nekko (driver); us (et-who, et-lab-start)

What a user sees

A card that has dropped off the PCIe bus looks present: /dev/et0_mgmt and /dev/et0_ops are there, the driver is bound, et-who and et-who --check report it free, and on 2 Oct et-lab-start also listed it as free. The kernel log shows a burst of corrected receiver errors, then only SQ[0] sync: head mismatched, head_remote: -1 and SQ[0] corrupt: circbuffer header invalid!. The only clear sign is /sys/bus/pci/devices/<card>/current_link_speed, which reads Unknown, with a width of 63.

What went wrong, and why
  • inferred The driver reads a dead link's all-ones as data and reports a corrupt queue; it has no link-down or surprise-removal path.
  • known A plain systemctl reboot at 10:47 on 2 Oct left the slot empty, because a warm reboot keeps the slot powered. Two reboots that cut the slot's power, with the kernel's reboot mode set to cold and its type to PCI as root, brought the card back, at 06:45 and 10:53.
How we found it

Watching aifoundry2's card fail and recover on 1–2 Oct (H28), and checking what each tool said about it.

What it cost

Time: a newcomer's check said the card was free, and recovering it took two reboots. A user could also start work on a card that is not there.

Where it stands

Before using a card, check that its link speed is not Unknown; our onboarding brief has done so since 2 Oct, and section 4.6 now does too. Recover a card only with a power-cutting reboot, as root, and only once the cause is fixed: for aifoundry2 that is its cooling (SH5). Our et-who and et-lab-start should check the link (H34).

What the Nekko team can do
  • Driver: detect a link loss or all-ones reads, log “card lost” plainly, mark the card dead and return ENODEV to callers (CF5).

Requests: CF5, SH5

Evidence
  • docs/findings/14-card-behaviour.md (the drops and the reboots, 2 Oct)
  • aifoundry2, read 4 Oct: current_link_speed Unknown, width 63; /dev/et0_* present; dmesg at 12:01:58 (root port: rollover, timeout) and 12:02:01 (SQ[0], head_remote: -1)

Back to the table · Requests

C33 Nothing reads the cards' DRAM temperature, and DRAM refresh never speeds up on a hot card#

New corrupts results latent all four cards

4 October New Since structural; found 28 Sep (E53), from the source

Found in the source and in one measurement. The memory set-up leaves the controllers' temperature derating commented out (“not needed for bring-up”), nothing reads the LPDDR4X's own temperature register (MR4), and the memory shires have no sensor. In E53 (28 Sep) refresh stayed at its programmed rate at die means of 52–75 °C. LPDDR4X needs faster refresh above 85 °C, and the dies have run at 90–138 °C. No wrong result was seen up to an 81 °C mean; nothing hotter was checked.

Next: Nekko team CF15

Who acts: Nekko (firmware, hardware)

What a user sees

Nothing, and that is the problem: no tool reports the memory's temperature, and nothing changes when the card is hot.

What went wrong, and why
  • known The DDR initialisation leaves the memory controllers' temperature derating commented out: “FUTURE derate is not needed for bring-up” (etsoc-hal memshire_ddr_init_functions.c lines 4686–4688 at 836a4ab). Nothing reads the LPDDR4X's own temperature register (MR4), and the memory shires carry no temperature sensor.
  • known Measured in E53 (28 Sep): refresh stayed at its programmed 1× (3,616 controller clocks at 933 MHz) at die means of 52–75 °C.
  • known LPDDR4X needs 2× or 4× refresh above 85 °C case temperature. The dies have run far hotter: aifoundry2 at a 90–103 °C mean on 26 Sep (C22), card 0 at a 119 °C mean in September (C21), and aifoundry2 at 138 °C on 2 Oct (H28).
  • unknown The DRAM's own temperature at those times, its part and its temperature grade.
How we found it

Our overheating study (E53, 28 Sep) asked whether a hot card loses data, read the memory set-up in the source, and checked 663 launches for wrong results.

What it cost

Nothing observed: all 663 checked launches, up to an 81 °C mean, were correct. It is a risk to any result taken on a hotter card.

Where it stands

Treat results from runs above about an 85 °C die mean as unchecked. The lab's 90 °C stop rule keeps most runs below that.

What the Nekko team can do
  • Name the DRAM part and its temperature grade; say whether derating was left off on purpose; in firmware, read MR4 and enable derating, or at least report the DRAM's temperature (CF15).

Requests: CF15, CF1

Evidence
  • docs/findings/05-claims.md (the row “DRAM refresh stays at 1× hot”, E53)
  • etsoc-hal memshire_ddr_init_functions.c lines 4686–4688 at et-platform 836a4ab

Back to the table · Requests

3.2 The host machines and access#

The three Ubuntu hosts: logging in, disks, power, boot, packages, logging, and the services that share the cards. 37 problems on 4 October: 14 fixed, 5 partly fixed, 11 open, 7 new; 8 of the open and partly fixed ones updated since 27 September.

H1 Tailscale SSH asks for a browser check, and the check link can 404#

Open workaround blocks work all three

4 October Open workaround Since 18 Sep

Unchanged. The check-mode steps and the 404 fix are in our repository's docs/lab-access.md and in Appendix A, but not yet on the public New user page that newcomers now follow (ours to add, AS3).

27 Sep: Unchanged; the onboarding text now exists in our repository and in the draft onboarding page.

Next: Nekko team AS3, DI1

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko / tailnet admin: policy and onboarding text

What a user sees

ssh aifoundry2 stops and prints a https://login.tailscale.com/a/… URL. Nothing happens until someone approves it in a browser while that ssh is still waiting, and the approval lapses after the tailnet's check period (12 h by default), so scripted or BatchMode ssh, and agents, just hang. Opening the URL sometimes gives Error 404 … could not be located.

What went wrong, and why

known The tailnet's SSH policy uses the "check" action for member logins, which needs a browser re-authentication with the member account every check period. inferred The 404 appears when the browser's login.tailscale.com session belongs to another account or tailnet than the one that issued the URL; logging out and back in cleared it.

How we found it

The owner hit the 404 while logging in on 25 Sep.

What it cost

A new user cannot log in at all until they find the log-out/log-in trick; an agent's ssh hangs indefinitely.

The workaround

Worked around by the owner on 25 Sep: log out of login.tailscale.com, log back in with the lab member account, run ssh again for a fresh URL, and approve it while ssh waits. Our docs do not have this yet (D18). Machine-to-machine ssh between the lab machines gets no check (H2), which is why our unattended automation runs from a lab machine.

What the Nekko team can do
  • Put the check-mode steps in the lab's onboarding page and invite email: which account, which tailnet, approve while ssh waits, how long it lasts, and the 404 fix.
  • Consider a longer check period for the member-to-lab-machine rule, or accept with a device-posture condition, so an overnight job does not need a browser.

Requests: AS3, DI1

Evidence

Re-check, 27 September

  • not testable from a lab machine (machine-to-machine ssh gets no check); tailscale 1.102.4; docs/lab-access.md has the check-mode steps and the 404 fix

Back to the table · Requests

H2 Any account on one lab machine can log in as root on the others#

Open security all three

4 October Open Since found 25 Sep

Unchanged on 1 Oct. Since 2 Oct the lab's own onboarding uses the shared root login on purpose: each newcomer logs in as root once to create their own account (the lab's message to newcomers, and our New user page), so any narrowing must keep a way to create accounts (AS1, AS5).

27 Sep: Unchanged: our ordinary account opened root shells on aifoundry1 and aifoundry3 again for the 27 Sep read-only re-check, with no password or check.

Next: Nekko team AS1, AS5

The 25 September write-up, kept as history:

Who acts (25 Sep): Roman / tailnet admin

What a user sees

From an ordinary account on aifoundry2, a root login to aifoundry1 or aifoundry3 succeeds with no password and no browser check.

What went wrong, and why

known All three machines carry the same Tailscale tag, and the tailnet's SSH policy lets that tag log in as root on the others. Tailscale identifies a tagged machine by its tag, not by the local account that runs ssh, as tailscaled's records of our root sessions show. inferred Not tested with another account: every local account on any of the three (about 30) gets the same root access. The documented way to add lab users (scripts/add-lab-user.sh, run from a machine with root access to the others) needs root ssh to each machine, which this path provides.

How we found it

Reading tailscaled's process line during the aifoundry1 fix, and the aifoundry2 root audit.

What it cost

Any account on the three machines can alter or delete anything on all of them, including other people's data and running measurements, and the logs show only a tag, not a person.

Where it stands

Nothing changed. We used exactly this path, at the owner's request, for all of the 25 Sep fixes.

What the Nekko team can do
  • Decide whether machine-to-machine root is needed. If it only serves admin scripts such as add-lab-user.sh, restrict it: root only for an admin group from their own devices (with check mode), and tag:sf to tag:sf limited to autogroup:nonroot.
  • Document the root paths that remain.

Requests: AS1

Evidence
  • aif1/fix-transcript.txt 15:00:53 (tailscaled be-child ssh … --local-user=root … --remote-user=tag:sf)
  • labfix/audit-aifoundry2.md S1 (root SSH is allowed between tag:sf lab machines)
  • tailscale status --json on aifoundry2, 25 Sep 17:2x (all three tagged tag:sf)
  • docs/lab-access.md (Adding someone, step 3)

Re-check, 27 September

  • 27 Sep 21:26-21:31: our ordinary account on aifoundry2 opened root shells on aifoundry1 and aifoundry3 with no password or check (the read-only re-check); 21 root sessions on aifoundry3 since 25 Sep, all ours

Back to the table · Requests

H3 aifoundry1's disk is full of user data: 95% of a single 452 GB pool#

Fixed 30 Sep updated blocks work aifoundry1

4 October Fixed 30 Sep updated Since 30 Sep 22:25 (a departed user's public model checkpoints deleted)

Fixed on 30 Sep at 22:25 PDT, at the owner's word: we deleted a departed user's public model checkpoints, the 27 files of 1 GB or more (118.1 GB of public models such as Qwen, Llama, Gemma, SmolVLM, RWKV, LFM and TinyLlama, which can be downloaded again). The account no longer existed, and nothing used the files (no process had them open, and no system setting or scheduled job named them); the smaller checkpoints, all code and a 10 GB compiled bundle were kept. /home went from 99% used (7.3 GB free) to 72% (116 GB free), and the pool from 95% to 71% (130 GB free of 452 GB). The other owners' data is unchanged (MO1); with no quotas the pool can fill again (PO4). Still so on 4 Oct: 115 GB free on /home (72%), and the pool 71% full with 128 GB free.

Our 25 Sep ZFS snapshots held about 2.4 GB: usedbysnapshots at 23:20 on 27 Sep was 2.14 GB across rpool's datasets and 0.29 GB on bpool (the per-snapshot USED column, 1.86 GB and 105 MB, undercounts the blocks snapshots A and B share); they held the files the 25 Sep upgrade replaced. On 28 Sep at 07:53 the dry run of their destroy (U3) listed 34 snapshots on the pool, all ours (17 each of @labfix-20260925-A and -B), with no holds; the destroy itself, refused by this session's permission check at 07:54, ran at 08:33:40 with the owner's approval. All 34 are gone: rpool's available space rose from 5.70 to 7.98 GB and bpool's from 1.36 to 1.65 GB. The pool was still 95% full then (fragmentation 59%). The user data, the Docker images and the three stopped containers were unchanged, so the rest of the space was the account owners' (MO1).

Next: Nekko team PO4 (so it does not fill again), MO1 (the rest, no longer urgent); us U18

The 25 September write-up, kept as history:

Who acts (25 Sep): Accounts rehan and roman; the lab (two homes whose accounts no longer exist, Docker, /root)

What a user sees

At 15:00 on 25 Sep, /, /home, /root and /var each had 154 MB free. Builds, logs and apt fail with No space left on device; the 24 Sep kernel update died half-way (H5).

What went wrong, and why

known One ZFS pool (rpool, 452 GB, no redundancy) holds the system and every home. 392 GB of it is /home, mostly GGUF models: /home/rehan 176 GB, a 135 GB home with no account, /home/roman 66 GB, an 11 GB home with no account, plus 9.8 GB of GGUF copies in /root/et-jobs-deploy and 7.8 GB of Docker images no container uses. There are no quotas. ZFS holds back about 3% of a pool, so every dataset reaches 0 bytes free near 96%. The PCIe log flood (C16) added about 100 MB a day.

How we found it

The aifoundry1 troubleshooting session (13:57: / had 164 MB left), then the fix and the root audit.

What it cost

The 24 Sep kernel update failed; any build or model download on aifoundry1 fails; one more 20 GB model fills it again.

Where it stands

Caches only, by us as root at the owner's request. 15:03 and 15:11: the apt cache (281 MB), archived journal (118 MB), ten disabled snap revisions, root's Hugging Face chunk cache (5.4 GB) and root's ccache (0.5 GB). 16:24–16:25 (A2, A14a): kernel 7.0.0-30, headers of kernels that are not installed, openipmi and old GitHub runner versions, about 1.9 GB. At 17:19 the pool is still 95% allocated (432 of 452 GB) with 5.94 GB available; our labfix-20260925-A/-B ZFS snapshots pin about 1 GB until they are destroyed (planned after 27 Sep). No user file was touched: the user data is for its owners.

What the Nekko team can do
  • The owners delete or move what they no longer need: account rehan (123 GB of models), account roman (one 52 GB directory untouched since October 2025; the fix log below names it); the lab for the two homes with no account, /root/et-jobs-deploy and the unused Docker images (docker container prune -f && docker image prune -a -f, after our snapshots are destroyed).
  • Then per-user quotas (zfs set userquota@<user>=100G rpool/USERDATA/home_2t5qyu), one shared read-only /models dataset, and an alert at 85%.

Requests: MO1, PO4

Evidence

Re-check, 27 September

  • aifoundry1 27 Sep: pool 452 G, CAP 95%, AVAIL 5.01 G (5.94 G at 25 Sep 17:19; 6.4 G after the cleanup), FRAG 51% -> 60%
  • our labfix-20260925-A/-B snapshots: usedbysnapshots 2,138,869,760 B on rpool and 289,382,400 B on bpool (23:20); files replaced by the 25 Sep upgrade
  • user data unchanged (the largest accounts and root-owned homes as in the report); Docker images and 3 stopped containers unchanged

Root steps, 28 September (verified 08:00–08:02; the steps after that, by their own output)

  • aifoundry1, 28 Sep 08:00 (verification): all 34 of our snapshots still there, and no other snapshot on the host; usedbysnapshots 1.93 GB on rpool/ROOT/ubuntu_jleyxm and 276 MB on bpool; et-lab-health warned on the pool and /home
  • 08:33:40 (U3): zfs destroy -rv of the four snapshot roots reclaimed 832 MB, 1.31 GB, 105 MB and 171 MB; afterwards 0 labfix snapshots; available space: rpool 5,698,203,648 → 7,977,537,536 B, bpool 1,356,611,584 → 1,645,809,664 B; zpool list: rpool 95% full, fragmentation 59% (labreport2/fix28/a1/s01-real.out)

Back to the table · Requests

H4 aifoundry1: ZFS permanent errors in four files, on a single disk with no backups#

Open corrupts results aifoundry1

4 October Open Since 13 Sep scrub

Unchanged: the 13 Sep scrub's 42 errors, no scrub since, no snapshot and no backup. The automatic scrub runs at 00:24 on Sunday 11 Oct. aifoundry1 was most likely switched off without a clean shutdown at about 12:55 on 2 Oct for the fan work (its journal stops mid-stream), another hard stop for this single disk.

27–28 Sep: No new errors. Since our snapshots went (28 Sep 08:33, U3), zpool status -v lists four files and nothing else: two model files and a git object in other users' homes, and a toolchain build's cc1plus. They are unchanged and wait for their owners (MO2); the scrub after that is ours (U17).

Next: Nekko team MO2, PO4; us U17

The 25 September write-up, kept as history:

Who acts (25 Sep): Account rehan; the owner of another home; the lab (build tree, backups); us (snapshots)

What a user sees

zpool status -v rpool says Permanent errors have been detected in the following files (the 13 Sep scrub found 42 errors). Reading those files returns EIO: git fsck fails in one tree, and two llama.cpp vocabulary files cannot be read.

What went wrong, and why

known A single-disk pool with no redundancy, no snapshots before 25 Sep, and no backups. inferred The frequent power losses (H7) and non-ECC memory; the disk shows 0 read, write and checksum errors since import, and SMART shows 0 media errors.

How we found it

The root audit of aifoundry1 (zpool status -v).

What it cost

Four files already lost; 392 GB of home directories has no copy.

Where it stands

Nothing restored or deleted (user files). Our labfix-20260925-A and -B snapshots also captured the corrupt cc1plus, so at 17:19 zpool status lists it three times; the list cannot clear until those snapshots are destroyed as well.

What the Nekko team can do
  • The owners restore or delete: account rehan (two vocabulary .gguf files under /home/rehan/auto/), the owner of another home (one git object: re-clone), the lab for cc1plus (delete the 4.9 GB /usr/src/riscv-gnu-toolchain/build-* trees, which /opt/et does not use).
  • We destroy the labfix snapshots; then, while idle: zpool scrub rpool && zpool wait -t scrub rpool && zpool status -v rpool && zpool clear rpool.
  • Decide a backup: zfs send of rpool/USERDATA to another machine, or at least scheduled snapshots once there is room.

Requests: MO2, PO4

Evidence

Re-check, 27 September

  • zpool status -v: 42 errors from the 13 Sep scrub; list = 3 user files + /usr/src/riscv-gnu-toolchain/.../cc1plus + the same cc1plus in our two snapshots; vdev 0/0/0; SMART clean (0 media errors, spare 100%, used 17%)
  • next automatic scrub Sun 11 Oct 00:24

Root steps, 28 September (verified 08:00–08:02; the steps after that, by their own output)

  • aifoundry1, 28 Sep 08:33:40, after U3: zpool status -v rpool lists 4 files with permanent errors: account rehan's two .gguf files, one git object in another user's home, and /usr/src/riscv-gnu-toolchain/build-gcc-newlib-stage1/gcc/cc1plus; none in a snapshot

Back to the table · Requests

H5 aifoundry1's 24 September kernel update stopped half-way#

Fixed 25 Sep blocks work aifoundry1

27 September Fixed 25 Sep Since 25 Sep 15:03 (us)

Holds: dpkg --audit is empty, and unattended upgrades ran cleanly on 26 and 27 Sep.

Next: Us U18

The 25 September write-up, kept as history:

Who acts (25 Sep): Done. Lab: monitor dpkg

What a user sees

The machine asked for a restart into 7.0.0-34, but that kernel was half-installed: no /boot/initrd.img-7.0.0-34-generic, GRUB not updated, and 10 pending entries in /var/lib/dpkg/updates.

What went wrong, and why

known unattended-upgrades ran out of space (H3) in the middle of the linux-image-7.0.0-34 postinst.

How we found it

The aifoundry1 troubleshooting session found no initramfs for the next kernel.

What it cost

inferred Without an initramfs, 7.0.0-34 cannot mount the ZFS root, so the next power cut would have left aifoundry1 down until someone came on site with a console.

What was done

15:03:40: dpkg --configure -a (exit 0) after freeing space. It ran DKMS for 7.0.0-34, wrote the initramfs (ZFS included) and ran update-grub (default entry 7.0.0-34). 7.0.0-31 stays installed as the fallback.

How to verify

15:03:58: linux-image-7.0.0-34 and linux-modules-7.0.0-34 read ii, /var/lib/dpkg/updates is empty, apt-get check is clean. 17:19: dpkg --audit prints nothing, and the initramfs rewritten by the 16:39 upgrade still holds the ZFS modules. Booting 7.0.0-34 on aifoundry1 is still untested (H6).

What the Nekko team can do
  • Monitor dpkg --audit (non-empty means a stuck update) and free space; keep unattended-upgrades off a pool above 90%.
Evidence

Re-check, 27 September

  • dpkg --audit empty; apt-get check ok; unattended-upgrades ran cleanly 26 and 27 Sep

Back to the table · Requests

H6 The reboot into kernel 7.0.0-34 has been pending since 24 Sep; done only on aifoundry3#

Fixed 30 Sep wastes time aifoundry1, aifoundry2

30 September Fixed 30 Sep Since 30 Sep 15:07 (the power cycle)

All three hosts run 7.0.0-34: aifoundry3 since 25 Sep, and aifoundry1 and aifoundry2 since Roman power-cycled the lab at about 15:07 on 30 Sep, after our link test hung aifoundry1 (C28). Section 4.6's checks passed on all three at 15:18–15:20.

  • aifoundry3 fixed: 7.0.0-34; nothing pending
  • aifoundry1 fixed 30 Sep 15:07: 7.0.0-34 after the power cycle (C28); driver loaded at boot
  • aifoundry2 fixed 30 Sep 15:07: 7.0.0-34 after the power cycle; nothing pending

Next: Nekko team PO3 (a maintenance window for the next one)

The 25 September write-up, kept as history:

Who acts (25 Sep): Lab: a window. aifoundry1: the logged-in user's OK. aifoundry2: us

What a user sees

All three machines have said "System restart required" since 24 Sep. Nobody knew whether 7.0.0-34 would boot, and a reboot kills every running job.

What went wrong, and why

known unattended-upgrades installs kernels, but reboots are manual, and nobody schedules them on shared, remote-only, Wi-Fi-only machines with no console (H21).

How we found it

The root audits (/var/run/reboot-required on all three).

What it cost

Until then two machines run a kernel other than the one they will boot after the next power cut, and host-side timing has to be re-baselined after it.

Where it stands

aifoundry3: rebooted by us at 16:20:36 (a clean shutdown record) and up at 16:21:05 on 7.0.0-34. At 17:19: et_soc1 0.20.0/47D26A30…, the clock guard pinned the card on attempt 1 with the current boot_id in its marker, and chrony, the core-dump pattern, the performance profile and Wi-Fi power saving off all survived the boot. aifoundry1 and aifoundry2: still pending at 17:19. aifoundry1 needs the OK of the other user who has been logged in since 18 Sep (the reboot ends their session) and someone reachable on site, because its /boot/grub is on ZFS and a one-time grub-reboot fallback cannot work there. aifoundry2's reboot ends our own session and wipes /tmp (H22).

What the Nekko team can do
  • Agree a standing maintenance window (weekly, say), announce it with wall and in the motd, and reboot all three in it. Keep 7.0.0-31 installed until 7.0.0-34 has booted everywhere; the post-boot checks are in section 4 below.

Requests: PO3

Evidence
  • read-only checks 17:19 on all three (kernel, reboot-required packages, clock-guard marker, last shutdown record)
  • labfix/hostlogs/aifoundry3/W3-post.log; fix plan B4, W1, W2

Re-check, 27 September

  • uname -r on aifoundry1 and aifoundry2: 7.0.0-31-generic

Back to the table · Requests

H7 Most unclean resets hit all three machines at once: their power is cut together#

Open corrupts results all three

4 October Open Since Jul (clusters 17 and 23 Jul); last 18 Sep

No unplanned cut since 18 Sep. All three restarted together, without shutdown records, at 15:07 on 30 Sep: Roman's deliberate power cycle after C28. aifoundry1 was most likely switched off without a clean shutdown at about 12:55 on 2 Oct for the fan work. A single machine should be powered off by itself, after sudo poweroff (SH2).

27–28 Sep: No new cut since 18 Sep; cause unknown. aifoundry3's wtmp, saved on 27 Sep, shows 114 of 125 boots since 2 Jan without a clean shutdown record; its journal, with the July boots, was exported on 28 Sep (H32).

wtmp on aifoundry3, saved as a user on 27 Sep at 22:17 (U9): 125 boots since 2 Jan and only 10 clean shutdown records. The boots without one cluster on 17 Jul (13), 23 Jul (11), 17 Apr (7) and 5 Mar (6). A missing shutdown record can have other causes (a hang or a hard reset), so this is a lead, not a conclusion; the journal would confirm the causes. On 28 Sep we exported it as root (48 boots back to 17 Jul 01:14) and saved each boot's last kernel and PID-1 lines (U9, H32), so the July record is kept for that.

Next: Nekko team SH2

The 25 September write-up, kept as history:

Who acts (25 Sep): Roman / on site

What a user sees

Machines come back with every job gone and no log of why. Since 24 May, aifoundry1 has 70 boots and 2 clean shutdown records; aifoundry2 77 and 10; aifoundry3 75 and 7 (one of them our reboot on 25 Sep).

What went wrong, and why

known 64 of the 85 boot events since 24 May hit all three machines within 3 minutes of each other (for example 18 Sep 02:27 and 15:43, ten times on 17 Jul and eleven times on 23 Jul). The NVMe drives count 11,276, 815 and 3,079 unsafe shutdowns over their lifetime (aifoundry1, 2, 3). inferred The power of all three is cut together, upstream of the machines (a shared strip, PDU or circuit), so these are not independent host hangs. unknown Whether the cuts are outages or someone cycling the shared supply: the two densest clusters, 17 and 23 Jul, fall on the days aifoundry3's runtime and clock guard were being changed (C18, C5), and on 17 Jul the lab admin asked users to leave hung cards to him for a power cycle (C14), which inferred suggests deliberate power cycles. Either way it corrects our first reading of aifoundry1's resets as something specific to that machine.

How we found it

The aifoundry1 audit counted the unclean boots; comparing wtmp across the three machines at the end of the day showed that they coincide.

What it cost

Each event kills every run on all three machines at once and re-initialises the cards; inferred the likely cause of aifoundry1's ZFS corruption (H4).

Where it stands

Nothing can be fixed remotely. The persistent journal (H9) will keep the logs of the next events.

What the Nekko team can do
  • Find out what cuts the power to all three at once (building outages, a breaker, a shared strip or PDU, or someone cycling it). If hung cards are recovered by cycling a shared supply, use the per-card reset or a per-machine switch instead, so one card's hang does not reset everyone (C14).
  • Put the three machines on a UPS with NUT so they shut down cleanly, and keep a log of on-site power work. The BIOS already powers on after AC loss, which is right for a remote lab.

Requests: SH2

Evidence
  • labreport/boots-clusters.txt (last -x on all three, read 17:2x)
  • read-only smartctl 17:19 (unsafe shutdowns); labfix/audit-aifoundry1.md F7

Re-check, 27 September

  • no power event on any host since 18 Sep (aifoundry3: only our clean 25 Sep reboot); unsafe-shutdown counters unchanged (aifoundry1 11,276 of 12,089 power cycles)
  • aifoundry3's July boot records, the evidence for the cuts, are about to be vacuumed (H32)

Root steps, 28 September (verified 08:00–08:02; the steps after that, by their own output)

  • aifoundry3, 28 Sep 06:44–06:50: the journal export (48 boots from 17 Jul 01:14) and 100 files of per-boot kernel and PID-1 tails (H32)

Back to the table · Requests

H8 The boot never finishes: the splash screen waits forever on headless machines#

Fixed 30 Sep workaround cosmetic all three

30 September Fixed 30 Sep workaround Since 30 Sep 15:07 (all three booted headless)

Fixed: all three hosts boot to the text target (U27), and after the power cycle of 30 Sep at about 15:07 all three reported running with no queued jobs (15:18–15:20).

Next: nothing left

The 25 September write-up, kept as history:

Who acts (25 Sep): Roman: the headless decision

What a user sees

systemctl is-system-running says starting for days and systemd-analyze says Bootup is not yet finished. systemctl is-system-running --wait, anything ordered after multi-user.target, and the text gettys never start.

What went wrong, and why

known plymouth-quit-wait.service never finishes, because gdm, which is supposed to end the boot splash, never does so on these machines with no monitor attached. A trap: systemctl start plymouth-quit.service would stop gdm (gdm has Conflicts=plymouth-quit.service); only the plymouth quit command is safe.

How we found it

The root audits (systemctl list-jobs).

What it cost

Monitoring reads "starting"; scripts that wait for boot completion hang.

The workaround

By us: plymouth quit at 16:15 (aifoundry3), 16:24 (aifoundry2) and 16:26 (aifoundry1), and again at 16:22:16 after aifoundry3's reboot, where the hang came back. At 17:19 all three say running with no queued jobs. It returns at every boot while the default target is graphical.target.

What the Nekko team can do
  • If nobody uses a monitor at the lab: systemctl set-default multi-user.target on all three before the next reboot (rollback: systemctl set-default graphical.target). That also removes about 42 greeter processes per machine.
Evidence
  • labfix/audit-aifoundry1.md F5; labfix/audit-aifoundry2.md O2; labfix/audit-aifoundry3.md F14
  • labfix/hostlogs/aifoundry3/W3-post.log (16:22:16 plymouth quit)

Re-check, 27 September

  • all three: is-system-running = running, no jobs; get-default = graphical.target; GRUB "quiet splash" (aifoundry2); ~43 gdm processes on aifoundry3

Back to the table · Requests

H9 aifoundry1's journal kept only about a day#

Fixed 25 Sep wastes time aifoundry1 (aifoundry2, aifoundry3 capped too)

4 October Fixed 25 Sep Since 25 Sep 16:15-16:24 (us, A9)

aifoundry1's journal is at its 1 GB cap again and has grown by about 170–185 MB a day since 2 Oct, mostly sudo lines from our own live monitor, one a second (H35). It still keeps about ten days, about five at this rate. aifoundry2's is at 1.14 GB of its 2 GB cap.

30 Sep: Holds on all three. On aifoundry3 our 2 GB cap filled and began deleting the July boots (H32); we raised it to 4 GB on 28 Sep (U9). aifoundry2's journal is at 1.24 GB of the same 2 GB cap (30 Sep), so it will fill within a month or two; raising it to 4 GB there too is a root step of ours.

  • aifoundry1 fixed: persistent, 1 G cap, 69 MB used; history since 24 Sep 15:00; keeping earlier boots is untested until its reboot
  • aifoundry3 fixed (cap raised 28 Sep): our cap was full at 2 G (1.8–1.9 G) and journald had begun deleting the July boots (H32); since 28 Sep 07:11 it is 4 G (SystemMaxUse=4G), with 1.8 G used and 46 boots kept, back to 17 Jul 01:21

Next: us H35 (stop the per-second sudo); aifoundry2's cap (root)

The 25 September write-up, kept as history:

Who acts (25 Sep): Done

What a user sees

journalctl -b -1 gave No journal boot entry found: only the current day was kept, so the history of the resets and of the 18 Sep PCIe error jump was gone.

What went wrong, and why

known journald's default SystemKeepFree (15% of the file system) left almost no room on the nearly full pool, and the AER flood (C16) filled what there was.

How we found it

The root audit of aifoundry1.

What it cost

The causes of the resets and of card 0's link degradation on 18 Sep cannot be reconstructed.

What was done

Fix A9, 16:24:34: /etc/systemd/journald.conf.d/60-labfix.conf with Storage=persistent, SystemMaxUse=1G, SystemKeepFree=1G, then a journald restart. aifoundry2 and aifoundry3 got SystemMaxUse=2G. If aifoundry1's pool drops below 1 GB free, journald shrinks again (H3).

How to verify

The VERIFY block shows the settings; the journal was 51 MB at 17:19. After aifoundry1's next boot, journalctl --list-boots should list two or more boots.

How to roll back

Remove the drop-in; systemctl restart systemd-journald.

What the Nekko team can do
  • Nothing beyond checking the next boot.
Evidence
  • labfix/audit-aifoundry1.md F6
  • labfix/hostlogs/aifoundry1/A-session-20260925T162431.log (A9, VERIFY)

Re-check, 27 September

  • journald.conf.d/60-labfix.conf in place on all three; aifoundry2 22 boots kept, aifoundry3 48 boots back to 17 Jul

Back to the table · Requests

H10 Time came from one NTP server over Wi-Fi#

Fixed 25 Sep corrupts results all three

27 September Fixed 25 Sep Since 25 Sep 16:14-16:26 (us, A6)

Holds: chrony is in sync with 8 sources on all three (offsets 0.09–0.33 ms).

Next: nothing left

The 25 September write-up, kept as history:

Who acts (25 Sep): Done

What a user sees

Timelines from different machines disagreed by 20–30 ms.

What went wrong, and why

known Ubuntu's default systemd-timesyncd with a single ntp.ubuntu.com server, polled every 34 minutes over Wi-Fi (delay 91–237 ms, jitter 23–50 ms). Offsets read −14.7, −0.5 and +7.5 ms in the audits (aifoundry1, 2, 3; 15:17–15:45) and −7.0, +1.6 and +24.9 ms just before the change (16:14–16:24).

How we found it

The root audits; the fix's backups of timedatectl timesync-status.

What it cost

Cross-machine event alignment was only good to a few tens of milliseconds before 25 Sep; single-machine data is unaffected.

What was done

Fix A6 (aifoundry3 16:14:50, aifoundry2 16:24:03, aifoundry1 16:26:24): chrony 4.5 replaces systemd-timesyncd, with Ubuntu's pool sources.

How to verify

chronyc -n tracking shows Leap status : Normal; at 17:19 each machine had 8 sources and offsets of 0.2–0.9 ms, also after aifoundry3's reboot.

How to roll back

apt-get install systemd-timesyncd (it removes chrony).

What the Nekko team can do
  • Keep chrony; once the machines are wired (H11), optionally peer them with each other.
Evidence
  • labfix/audit-aifoundry1.md F11; labfix/audit-aifoundry2.md O3; labfix/audit-aifoundry3.md F8
  • labfix/hostlogs/*/timesync-*.txt (the A1 backups, offsets before the change)
  • read-only checks 17:19 (chronyc)

Re-check, 27 September

  • chrony Leap status Normal, 8 sources, offsets 0.09-0.33 ms on all three

Back to the table · Requests

H11 No Ethernet: all three machines run on Wi-Fi#

Partly fixed wastes time all three

30 September Partly fixed Since 25 Sep 16:15-16:26 (Wi-Fi power saving off, us)

The mitigation holds (Wi-Fi power saving off); still no cable, and aifoundry2's roaming got worse: it switched access points 45 times on 30 Sep by 14:47 (3–25 a day on 26–29 Sep).

Next: Nekko team SH4

The 25 September write-up, kept as history:

Who acts (25 Sep): On site: plug in the cables

What a user sees

Slow rsync and model transfers, SSH sessions that drop, a 12-minute Tailscale outage on aifoundry2 on 23 Sep, CI broker timeouts. If the Wi-Fi goes, every way in goes with it, root included.

What went wrong, and why

known enp7s0 (the same 10GbE port on all three) has no carrier on any of the three; all traffic goes over Wi-Fi, which had power saving on (227,796 missed beacons on aifoundry1, up to 51 access-point roams a day on aifoundry2).

How we found it

The root audits.

What it cost

Every transfer and remote session goes over a roaming Wi-Fi link; one outage cuts off all access.

Where it stands

Mitigated by us (A16, 16:15–16:26): Wi-Fi power saving off on all three (the NetworkManager drop-in zz-labfix-wifi-powersave.conf with wifi.powersave = 2, plus iw dev wlp8s0 set power_save off at runtime on aifoundry1 and aifoundry3). At 17:19 all three read Power save: off. Rollback: remove the drop-in; iw dev wlp8s0 set power_save on. The cable itself is on-site work: enp7s0 still has no carrier.

What the Nekko team can do
  • Plug Ethernet into enp7s0 on all three. NetworkManager's existing "Wired connection 1" comes up by itself, and Wi-Fi stays as the fallback.

Requests: SH4

Evidence
  • labfix/audit-aifoundry1.md F10; labfix/audit-aifoundry2.md N1; labfix/audit-aifoundry3.md F7
  • labfix/consistency.md D13 (the 23 Sep outage)
  • read-only checks 17:19 (iw, carrier)

Re-check, 27 September

  • enp7s0 carrier 0 on all three; power_save off holds
  • aifoundry3: missed-beacon lines 0 since the reboot, but it roams between two access points up to 16 times a day at -70 dBm; aifoundry2: 13 (26 Sep) and 25 (27 Sep) deauth/disconnects a day

Back to the table · Requests

H12 CI runners run as root and can take a card at any time#

Open corrupts results aifoundry1, aifoundry2 (aifoundry3's is dead)

27 September Open Since found 24-25 Sep

Unchanged: both runners run as root, enabled and idle (no job since 15 Jun and 24 Jul), still polling GitHub.

  • aifoundry3 open (dead runner): unit inactive/disabled since 17 Jul; /root/actions-runner-hf-hackathon still present: retire it

Next: Nekko team MO3, PO2

The 25 September write-up, kept as history:

Who acts (25 Sep): The CI owners; Roman

What a user sees

A push to the CI's GitHub repository can start a job that opens a card while someone is measuring (the nodes take one opener, so one side fails), and the job runs as root on a shared machine (one workflow runs apt-get install).

What went wrong, and why

known Self-hosted GitHub Actions runners are installed as root services: actions.runner.nekkoai-hf-hackathon.aifoundry1-et-soc1 (no User=, so root) and actions.runner.aifoundry-org-hf-hackathon.aifoundry2-et-soc1 (User=root), both active and enabled; their last jobs ran on 15 Jun and 24 Jul. aifoundry3's runner is disabled because its registration was deleted on GitHub on 17 Jul. The runner directories sit in /root, owned by individual accounts' UIDs. The CI repository's benchmark scripts do take flock on /var/lock/etsoc-shire0.lock, the same file as the lab's new card lock (/var/lock is /run/lock). unknown Whether every job that opens a card takes it; by default the scripts take card 0's lock, while on aifoundry1 the stock tools open both cards (C13).

How we found it

The aifoundry1 fix (the runner had to be stopped first) and the root audits.

What it cost

Nothing measured yet. A job during a measurement, if either side skips the lock, would crash it or be crashed, and a root job can change the machine under everyone. Our runners check for a running Runner.Worker before every block.

Where it stands

Nothing changed about the runners. Our queues treat a running Runner.Worker as "others present" and do not start, and they hold the same lock files; we stopped the runners on aifoundry1 and aifoundry2 for the fix windows, and they were restarted after the upgrade (17:19: active and enabled, no job running).

What the Nekko team can do
  • The CI owners decide per runner: disable it if the hackathon is over (systemctl disable --now actions.runner.<name>.service), or reinstall it under an unprivileged user (the cards are 0666); and make sure every job that opens a card holds /run/lock/etsoc-shire<N>.lock for the card it uses (the benchmark scripts already take card 0's).
  • Retire aifoundry3's dead runner; at least pause the runners during measurement campaigns.

Requests: MO3, PO2

Evidence
  • labfix/audit-aifoundry1.md F15; labfix/audit-aifoundry2.md E6 (benchmark-board.yml l.529, apt-get install); labfix/audit-aifoundry3.md F18
  • read-only check on aifoundry1 18:07: the runner's checkout (15 Jun) takes the lock in .github/ci/platform/deploy/soc3-benchmark.sh and .github/ci/scripts/run_llama_server_benchmark.py
  • labfix/consistency.md D1, D2; tools/claims-v3/lib.sh (Runner.Worker in OTHER_COMM)

Re-check, 27 September

  • aifoundry1 runner: root, enabled, active, no job since 15 Jun, polling daily; aifoundry2 runner: root, enabled, active, no job since 24 Jul, polling

Back to the table · Requests

H13 aifoundry3's demo web app can launch jobs on the card as root, outside any lock#

Open corrupts results aifoundry3

27 September Open Since found 25 Sep

Unchanged: the demo and the chatbot still run, with no requests since 25 Sep; disabling them no longer affects the driver (C2).

Next: Nekko team MO3, AS1

The 25 September write-up, kept as history:

Who acts (25 Sep): aifoundry3's admin (the demo's owner)

What a user sees

A web page on the tailnet (port 5000) queues model jobs that run on aifoundry3's card as root, possibly in the middle of someone's measurement.

What went wrong, and why

known etsoc1-demo.service (enabled, root, a Flask development server bound to the Tailscale address) has only an in-app lock and does not take the card lock. etsoc1-chatbot.service (root, 127.0.0.1:7860) runs while its backend is disabled and points into the home of an account that no longer exists. The unit files are owned by a non-root account. No POST requests in 30 days; last run 11 Jul.

How we found it

The aifoundry3 root audit.

What it cost

Low probability (idle for two months), but one launch would collide with any run on that card.

Where it stands

Not changed (someone else's service); at 17:19 both are active and enabled. Since C2 the demo is no longer needed to load the driver, so disabling it is now safe.

What the Nekko team can do
  • The demo's owner decides: disable both (systemctl disable --now etsoc1-demo.service etsoc1-chatbot.service), or have the demo take flock /run/lock/etsoc-shire0.lock around launches and bind to localhost or require a login.
  • chown root:root the unit files.

Requests: MO3, AS1

Evidence
  • labfix/audit-aifoundry3.md F4, F19; labfix/consistency.md D3
  • read-only check 17:19 (both services active, listening on :5000 and 127.0.0.1:7860)

Re-check, 27 September

  • etsoc1-demo and etsoc1-chatbot active/enabled (since the 25 Sep reboot), listening on aifoundry3's tailnet address (port 5000) and on localhost; 0 requests since 25 Sep; chatbot and llama-server unit files owned by a user account; the demo is no longer needed to load the driver (C2)

Back to the table · Requests

H14 numpy, venv, scipy and build packages were missing; aifoundry1 could not make a venv at all#

Fixed 25 Sep blocks work aifoundry1, aifoundry3 (aifoundry2 partly)

27 September Fixed 25 Sep Since 25 Sep 16:14-16:26 (us, A4)

Holds: numpy 1.26.4, scipy and venv on all three.

Next: nothing left

The 25 September write-up, kept as history:

Who acts (25 Sep): Done. Lab: one package list

What a user sees

import numpy failed on aifoundry1 and aifoundry3. On aifoundry1 python3 -m venv failed (no ensurepip) and pip refused ("externally managed", PEP 668), so a user could not get numpy without root. scipy, matplotlib and pandas were missing everywhere; aifoundry1 lacked pkg-config and about 74 build packages the other two had; aifoundry3 had no tmux, htop or smartctl.

What went wrong, and why

known The machines were set up by hand at different times; only aifoundry2 had python3-numpy; Ubuntu 24.04 blocks pip installs into the system Python.

How we found it

Our analysis scripts failed on aifoundry1 and aifoundry3; the audits listed the rest.

What it cost

Our analysis scripts failed on two machines, and every user had to build a private workaround (ours: a venv on aifoundry3, numpy unpacked into a home directory on aifoundry1).

What was done

Fix A4 (aifoundry3 16:14, aifoundry2 16:23, aifoundry1 16:25), with --no-install-recommends: python3-numpy 1.26.4 (as on aifoundry2), python3-venv and python3.12-venv, scipy, matplotlib, pandas, pkg-config and the build -dev libraries on aifoundry1, libfftw3-dev (gp-sdk), tmux on aifoundry3, htop, smartmontools, nvme-cli, lm-sensors, ripgrep and a few more.

How to verify

Each VERIFY block: numpy [1.26.4] 1.26.4, venv [ok] ok.

How to roll back

apt-get purge the same list (fix plan A4), then apt-get autoremove --purge; the package lists before the change are in /root/labfix-20260925/.

What the Nekko team can do
  • Keep one package list for lab machines and install it on any new machine; et-lab-manifest helps check that the hosts stay alike.
Evidence
  • labfix/audit-aifoundry1.md F3; labfix/audit-aifoundry3.md F5, F6; labfix/consistency.md C2–C5
  • labfix/dry-aifoundry1.txt (before: No module named numpy, No module named ensurepip)
  • labfix/hostlogs/ A-session VERIFY blocks

Re-check, 27 September

  • numpy 1.26.4, scipy 1.11.4, venv ok on all three

Back to the table · Requests

H15 Users could not read the kernel log, where the driver explains refused opens and card errors#

Fixed 25 Sep wastes time all three

27 September Fixed 25 Sep Since 25 Sep 16:22-16:28 (us, C1.1)

Holds. (Users still cannot read the system journal: by design.)

Next: Nekko team DI1

The 25 September write-up, kept as history:

Who acts (25 Sep): Done

What a user sees

dmesg said Operation not permitted, and users are not in adm or systemd-journal. aifoundry1's driver failure, the reason behind every EBUSY, 95 Minion Runtime Error events on aifoundry2 and aifoundry3's segfault lines were all invisible to the people who hit them; perf refused to run.

What went wrong, and why

known kernel.dmesg_restrict = 1 (the Ubuntu default) and kernel.perf_event_paranoid = 4 on all three.

How we found it

Our 22 Sep attempt on aifoundry1 could not read the driver's messages; the root audits confirmed the setting everywhere.

What it cost

It is part of why aifoundry1's driver failure took days to diagnose (C1).

What was done

Fix C1.1, approved by the owner (aifoundry3 16:22, aifoundry1 16:27, aifoundry2 16:28): /etc/sysctl.d/60-labfix-debug.conf sets kernel.dmesg_restrict = 0 and kernel.perf_event_paranoid = 2 (users may also profile their own processes); kptr_restrict stays 1. Users were deliberately not added to adm, which would expose everyone's SSH command lines (D15).

How to verify

dmesg | tail -1 as an ordinary user (works on all three at 17:18–17:19; on aifoundry1 it shows card 0's AER lines).

How to roll back

Remove the file; sysctl -w kernel.dmesg_restrict=1 kernel.perf_event_paranoid=4.

What the Nekko team can do
  • Keep it, and name dmesg in the onboarding page as the first place to look when a card will not open.

Requests: DI1

Evidence
  • labfix/audit-aifoundry2.md E4; labfix/audit-aifoundry3.md F10; labfix/audit-aifoundry1.md F12
  • labfix/hostlogs/aifoundry1/W1-noreboot.log, labfix/hostlogs/aifoundry2/W2-noreboot.log, labfix/hostlogs/aifoundry3/W3-post.log (C1.1)

Re-check, 27 September

  • dmesg_restrict 0, perf_event_paranoid 2 on all three; dmesg works as a user

Back to the table · Requests

H16 Crashing programs left no core dumps#

Fixed 25 Sep wastes time all three

4 October Fixed 25 Sep Since 25 Sep 16:14-16:26 (us, A7)

Proven in use. Side effect: apport's coredump hook also wrote .crash copies (C19) until it was switched off, on aifoundry1 on 28 Sep and on aifoundry2 and aifoundry3 on 30 Sep (U20); aifoundry2's failed hook unit cleared with its reboot; only apport's crash report of 28 Sep is left there (C19).

Next: Nekko team DI1

The 25 September write-up, kept as history:

Who acts (25 Sep): Done

What a user sees

aifoundry3's host programs segfault in about 1 launch in 100 (C17), and nothing was kept to debug it.

What went wrong, and why

known kernel.core_pattern piped cores to apport, which ignores binaries that do not belong to a package, and ulimit -c was 0.

How we found it

Chasing aifoundry3's intermittent crash.

What it cost

The aifoundry3 crash could not be debugged; each crash cost a rerun.

What was done

Fix A7 (aifoundry3 16:14:56, aifoundry2 16:24:09, aifoundry1 16:26:31): systemd-coredump installed (replacing apport-core-dump-handler), /etc/systemd/coredump.conf.d/60-labfix.conf (external storage, compressed, MaxUse 8G, 1G on aifoundry1) and /etc/security/limits.d/60-labfix-core.conf (soft core unlimited for login sessions). The new core_pattern hands systemd-coredump a fixed size limit instead of the process's own, so a core is kept even when ulimit -c is 0, as it still is for commands started over non-interactive Tailscale SSH (checked 18:07); no script has to raise it.

How to verify

17:47:48 on aifoundry3: our own enercat_host, started by a queue whose ulimit -c is 0, aborted and its core was stored (coredumpctl list: SIGABRT present; C17). sysctl -n kernel.core_pattern reads |/usr/lib/systemd/systemd-coredump %P %u … on all three at 17:19, also after aifoundry3's reboot. Users see their own cores with coredumpctl; root sees all.

How to roll back

apt-get install apport-core-dump-handler, then remove the two files.

What the Nekko team can do
  • Mention coredumpctl list and coredumpctl gdb <pid> in the onboarding page (they are in the motd).

Requests: DI1

Evidence
  • labfix/audit-aifoundry3.md F1 (apport: executable does not belong to a package, ignoring)
  • labfix/hostlogs/ A-session A7 and VERIFY blocks; read-only checks 17:19
  • read-only check on aifoundry3 18:07: core_pattern passes 9223372036854775808 as the limit; coredumpctl list (our process); the queue's own Max core file size soft limit 0

Re-check, 27 September

  • systemd-coredump in use: 10 cores on aifoundry1 (26 Sep), 2 on aifoundry2 (our sys_emu runs), 12 on aifoundry3 (they gave C17's root cause)

Back to the table · Requests

H17 Host CPUs ran in power-saving mode, so host-side timing jittered#

Fixed 25 Sep corrupts results all three

27 September Fixed 25 Sep Since 25 Sep 16:15-16:26 (us, A15)

Holds: the performance profile on all three.

Next: nothing left

The 25 September write-up, kept as history:

Who acts (25 Sep): Done. Users: record et-lab-manifest

What a user sees

Launch, sync and copy overheads and telemetry read times jittered with host frequency ramps, and host-side timings differ between machines (their CPUs and RAM differ too).

What went wrong, and why

known intel_pstate powersave with EPP balance_performance and power-profiles-daemon balanced on all three. The hosts differ: an i7-11700K with 128 GB (aifoundry1), an i5-11600 with 64 GB (aifoundry2), an i7-11700K with 32 GB (aifoundry3). Their runtime builds (C18) and, until the reboots, their kernels (H6) differ too.

How we found it

The root audits and the cross-machine consistency check.

What it cost

Host-side timing noise, and a one-time break in host-side comparability on 25 Sep.

What was done

Fix A15 (aifoundry3 16:15:07, aifoundry2 16:24:21, aifoundry1 16:26:44): powerprofilesctl set performance (persists across boots). Host-side timings from before about 16:15 on 25 Sep are not comparable with later ones; the hardware difference stays, and et-lab-manifest records it.

How to verify

powerprofilesctl get reads performance on all three at 17:19, also after aifoundry3's reboot.

How to roll back

powerprofilesctl set balanced.

What the Nekko team can do
  • Record et-lab-manifest with every measurement; compare host-side timings only within one machine; after the reboots, freeze the configuration for a measurement campaign.
Evidence
  • labfix/audit-aifoundry1.md F21; labfix/audit-aifoundry2.md O4; labfix/audit-aifoundry3.md F9; labfix/consistency.md B1, B3
  • labfix/hostlogs/ A-session A15 and VERIFY blocks

Re-check, 27 September

  • powerprofilesctl = performance, EPP performance on all three

Back to the table · Requests

H18 Months of pending updates, and a plain apt upgrade over Tailscale SSH kills itself#

Fixed 25 Sep wastes time all three

4 October Fixed 25 Sep Since 25 Sep 16:19-16:39 (us, B3)

Holds; recurs by design without a maintenance window. Docker 29 / containerd 2 are held on aifoundry1 only, for the CI owners (aifoundry2 and aifoundry3 already run Docker 29.1.3 and containerd 2.2.1). On 4 Oct no security update was pending on any host; 28, 21 and 16 others were, among them kernel 7.0.0-38 (PO3).

Next: Nekko team PO3

The 25 September write-up, kept as history:

Who acts (25 Sep): Done. CI owners: Docker on aifoundry1

What a user sees

213–222 non-security updates were pending on each machine (gdm, apparmor, snapd, fwupd, boost, binutils, tailscale). Upgrading tailscale restarts tailscaled, which drops the Tailscale SSH session that is running dpkg. aifoundry1's set included Docker 28→29, containerd 1.7→2.3 and compose 2→5.

What went wrong, and why

known unattended-upgrades is configured for noble-security only, and the tailscale package restarts its daemon on upgrade.

How we found it

The root audits.

What it cost

A naive upgrade over ssh would leave dpkg half-done when the session drops.

What was done

Fix B3: one full upgrade per machine, run detached with systemd-run --unit=labfix-upgrade … (aifoundry3 finished 16:19, aifoundry2 16:28–16:36, aifoundry1 16:27–16:39; 216, 212 and 200 packages upgraded). On aifoundry1 Docker was held and ZFS snapshots @labfix-20260925-B were taken first; the CI runners were stopped and restarted around it. The Docker and containerd major upgrade on aifoundry1 stays open for the CI owners to test.

How to verify

17:19: labfix-upgrade Result=success ExecMainStatus=0, dpkg --audit empty, tailscale 1.102.4, 14, 8 and 3 packages still upgradable (held Docker and phased updates).

How to roll back

Package downgrades are impractical (Ubuntu SRUs). On aifoundry1, roll the system datasets back to @labfix-20260925-B from a rescue shell; apt-mark unhold the Docker packages.

What the Nekko team can do
  • Run upgrades in the maintenance window (H6) through a detached unit; let the CI owners test Docker 29 and containerd 2 before unholding them.

Requests: PO3

Evidence
  • labfix/audit-aifoundry1.md F8, R6; labfix/audit-aifoundry2.md O1; labfix/audit-aifoundry3.md F12
  • labfix/hostlogs/*/B3-upgrade.log; labfix/hostlogs/aifoundry1/W1-noreboot.log (holds, snapshots)

Re-check, 27 September

  • dpkg --audit empty everywhere; upgradable: aifoundry1 14 (6 held Docker + 8 phased), aifoundry2 8 phased, aifoundry3 3
  • aifoundry3: noble-updates is not in the unattended-upgrades origins, so bug-fix updates pile up until someone runs a full upgrade

Back to the table · Requests

H19 Sudo, passwords and OpenSSH do not match the documented policy#

Partly fixed security aifoundry1, aifoundry2 (aifoundry3 partly)

4 October Partly fixed Since found 25 Sep

Sudo grew on 2 Oct: 7 members on aifoundry1, 18 on aifoundry2 and 3 on aifoundry3 (4 Oct). One new account got sudo on all three hosts that day, and a new file appeared in aifoundry2's /etc/sudoers.d; the New user page promises accounts with no sudo. Password logins stay off on aifoundry1 and aifoundry2. Our public accounts runbook still says to give sudo by setting a starting password (U7).

30 Sep: Password logins are off on both hosts with sshd: aifoundry1 since 28 Sep 20:51, aifoundry2 since 30 Sep 14:37 (U22); both offered only public-key logins at 14:40. The accounts and sudo stay with the lab (AS2), and the correction to our public accounts runbook waits for the owner's OK (U7).

  • aifoundry1 partly fixed 28 Sep 20:51: the login-service setting in place (U22; checked again at 01:10 on 29 Sep); sudo unchanged
  • aifoundry2 partly fixed 30 Sep 14:37: password logins off (U22, run by the owner); sudo 18 members on 4 Oct (17 on 30 Sep)
  • aifoundry3 open (partly): no sshd; 3 sudo members (4 Oct)

Next: Nekko team AS2; us U7, U22

The 25 September write-up, kept as history:

Who acts (25 Sep): Roman

What a user sees

docs/lab-access.md says accounts have no password and cannot sudo, but aifoundry2 has 17 sudo members, 16 of them with usable passwords (mostly hackathon-era accounts); one account has NOPASSWD sudo on aifoundry1 and aifoundry2 (added 18 Sep); OpenSSH on aifoundry1 and aifoundry2 accepts passwords from the Wi-Fi LAN; each machine has home directories with no matching account.

What went wrong, and why

known Accounts and sudo were granted ad hoc during events, and OpenSSH was left at Ubuntu's defaults next to Tailscale SSH.

How we found it

The root audits.

What it cost

No effect on work today; a guessed LAN password on aifoundry2 gives root.

Where it stands

Not changed (policy). At 17:19: sudo groups of 6, 17 and 2 members on aifoundry1, 2 and 3; sshd accepts passwords on aifoundry1 and aifoundry2; aifoundry3 has no sshd (H24).

What the Nekko team can do
  • Review sudo: gpasswd -d <user> sudo and passwd -l <user> for dormant accounts; review the NOPASSWD rule.
  • printf 'PasswordAuthentication no\n' > /etc/ssh/sshd_config.d/60-nopassword.conf && sshd -t && systemctl reload ssh (Tailscale SSH is unaffected).
  • Archive homes with no account (for d in /home/*; do id -u "${d##*/}" >/dev/null 2>&1 || echo "$d"; done lists them). We correct docs/lab-access.md (D18).

Requests: AS2

Evidence
  • labfix/audit-aifoundry2.md S1; labfix/audit-aifoundry1.md F22; labfix/consistency.md D14
  • docs/lab-access.md ("Accounts have no password, so they cannot use sudo")
  • read-only checks 17:19 (sshd settings, sudo group sizes)

Re-check, 27 September

  • aifoundry1 and aifoundry2: password logins over the LAN still accepted; the sudo members and the one NOPASSWD rule on each as on 25 Sep; nothing changed in passwd, shadow, group or sudoers since 25 Sep

Back to the table · Requests

H20 Host firmware: 2021 BIOS versions, and SSD firmware with a known health bug#

Open corrupts results latent all three

4 October Open Since found 25 Sep

Unchanged: BIOS F5, F5 and F6, and SSD firmware 3B2QGXA7 on all three. Since the 30 Sep boot only aifoundry1 has the IRQ 9 storm (aifoundry2's interrupt rate now matches aifoundry3's, so the storm comes and goes), and all three log the same ACPI BIOS error at boot (\ADBG, AE_ALREADY_EXISTS).

30 Sep: Unchanged, and wider: aifoundry1 has the same ACPI interrupt storm as aifoundry2 (IRQ 9 disabled after 100,001 interrupts and polled, about ten million SCI events each by 30 Sep), both on BIOS F5; aifoundry3, on F6, logs an ACPI error at boot instead.

Next: Nekko team SH3

The 25 September write-up, kept as history:

Who acts (25 Sep): On site (console)

What a user sees

Nothing visible yet from the SSDs. aifoundry2 disables IRQ 9 at boot (nobody cared, an ACPI interrupt storm) and polls ACPI instead; all three report gather_data_sampling: Vulnerable.

What went wrong, and why

known The boot drives are Samsung 980 PRO (500 GB, 1 TB, 500 GB) on firmware 3B2QGXA7, the release with Samsung's known health-degradation problem, fixed in 5B2QGXA7; their health counters are fine today (0 media errors; aifoundry1 17% used, spare 100%). The boards run Gigabyte Z590 AORUS MASTER BIOS F5 (June 2021) on aifoundry1 and aifoundry2 and F6 (August 2021) on aifoundry3, microcode 0x65.

How we found it

The root audits; smartctl on all three at the end of the day.

What it cost

Latent: a drive failure on aifoundry1 would lose data that has no backup (H4); host-side timing changes after a BIOS update.

Where it stands

Not changed: both need the console, and a BIOS update resets its settings.

What the Nekko team can do
  • At a console visit: back up first (H4), update the SSDs to 5B2QGXA7 (Samsung's tool or fwupd), enrol the module key (C20), then bring all three to the same current BIOS and check Secure Boot and "power on after AC loss" afterwards. Re-baseline host timing.

Requests: SH3

Evidence
  • read-only checks 17:19 (smartctl, dmidecode, /sys/devices/system/cpu/vulnerabilities)
  • labfix/audit-aifoundry2.md O7; labfix/audit-aifoundry1.md R5, F22; fix plan C3.4, C3.5

Re-check, 27 September

  • BIOS F5 06/18/2021 (aifoundry1, aifoundry2), F6 08/30/2021 (aifoundry3); gather_data_sampling Vulnerable; Samsung 980 PRO firmware 3B2QGXA7 on all three (health fine: spare 100%, used 5-17%)
  • aifoundry2: ACPI SCI interrupt storm, IRQ 9 disabled and polled (sci_not 7.76 M, ~9.6/s); aifoundry3: ACPI BIOS Error [\ADBG] at boot

Back to the table · Requests

H21 No console or out-of-band access, and the boot menu was hidden#

Partly fixed wastes time all three

30 September Partly fixed updated Since 25 Sep 16:15-16:28 (us, B2: menu)

The 5 s boot menu holds; still nobody at a console, and on 30 Sep it mattered: aifoundry1 went down and nothing could restart it remotely (C28, SH8).

Next: Nekko team SH6, SH8

The 25 September write-up, kept as history:

Who acts (25 Sep): Roman / on site

What a user sees

If a new kernel fails to boot, nobody remote can pick the old one: the GRUB menu was hidden with a 0 s timeout, there is no IPMI or KVM, and aifoundry1's /boot/grub is on ZFS, so a one-time grub-reboot fallback cannot work there.

What went wrong, and why

known Consumer boards (no BMC), Ubuntu's default hidden GRUB menu, and a ZFS root on aifoundry1.

How we found it

Planning the reboots.

What it cost

A failed boot means the machine is down until someone is on site.

Where it stands

The hidden menu was fixed by us (B2: aifoundry3 16:15:36, aifoundry1 16:27:43, aifoundry2 16:28:02): /etc/default/grub.d/60-labfix-menu.cfg with GRUB_TIMEOUT_STYLE=menu and GRUB_TIMEOUT=5, then update-grub. At 17:19 grub.cfg has set timeout=5 and set timeout_style=menu on all three. Rollback: remove the file; update-grub. Someone still has to be at the console to use it.

What the Nekko team can do
  • An IP-KVM (PiKVM or similar) on at least aifoundry1, or a named on-site contact for every maintenance window.

Requests: SH6

Evidence
  • labfix/audit-aifoundry1.md R1; labfix/audit-aifoundry2.md R2; fix plan B2, B4
  • labfix/hostlogs/aifoundry1/W1-noreboot.log, labfix/hostlogs/aifoundry2/W2-noreboot.log, labfix/hostlogs/aifoundry3/W3-config.log (B2)

Re-check, 27 September

  • GRUB set timeout_style=menu, timeout=5 on all three; no console or IP-KVM

Back to the table · Requests

H22 /tmp is wiped at boot, and coding agents keep their working files there#

Partly fixed updated corrupts results aifoundry2 (the same rule on all three)

4 October Partly fixed updated Since 27 Sep 22:12 (our files copied out); the risk grew 25–27 Sep

Our working files now live in the home directory (our rule since 30 Sep), so there is no /tmp copy to make, and U1 is closed. The rule itself is by design.

30 Sep: It happened: the power cycle of 30 Sep at about 15:07 cleared aifoundry2's /tmp, and 19 GB of our working files since 28 Sep 06:37, the last copy (U1), were lost; the repository and home directories were not touched. Keep work in the home directory: the banners, our docs and our public newcomer page say so.

  • aifoundry2 partly: /tmp/claude-1019 (19 GB) was cleared by the reboot of 30 Sep 15:07; the copy in ~/claude/private/ from 28 Sep 06:37 (15 GB) is what is left
  • aifoundry3 open (low impact): our files live in ~/nekko; /tmp holds 7.1 MB

The banner line “/tmp is cleared at every boot” is installed on all three hosts since 28 Sep (U5); docs/lab-access.md and the onboarding draft (appendix A) say it too.

Next: Nekko team DI1

The 25 September write-up, kept as history:

Who acts (25 Sep): Us (save before the reboot); the lab (a motd note)

What a user sees

Everything under /tmp is deleted at the next boot, and files untouched for 30 days are cleaned anyway. Our Claude Code scratchpad on aifoundry2 (/tmp/claude-1019, 5.2 GB: the audits, the fix plan, raw logs and validation data) would be lost.

What went wrong, and why

known /usr/lib/tmpfiles.d/tmp.conf has D /tmp 1777 root root 30d, and /tmp is on the root file system. With unannounced power losses (H7), the next boot can come at any time.

How we found it

The aifoundry2 audit's list of reboot risks.

What it cost

It would lose the day's raw evidence if the machine rebooted before it is copied.

Where it stands

aifoundry2's reboot is still pending (H6). Before it we copy the scratchpad to our home directory and commit the reports; the change logs also exist on each machine in /root/labfix-20260925/.

What the Nekko team can do
  • Say in the motd that /tmp is cleared at boot, so users and agents keep anything valuable in their home directory.

Requests: DI1

Evidence
  • labfix/audit-aifoundry2.md R1
  • read-only check on aifoundry2 17:2x (D /tmp 1777 root root 30d; du -sh /tmp/claude-1019 5.2G)

Re-check, 27 September

  • no banner on any host says /tmp is cleared at boot; aifoundry2's reboot is still pending (H6)

Root steps, 28 September (verified 08:00–08:02; the steps after that, by their own output)

  • all three, 28 Sep (aifoundry2 07:09, aifoundry3 07:26, aifoundry1 07:54): the installed banner (/etc/motd) equals this host's tools/lab/motd-<host> by sha256 (checked again at 08:00–08:02); the 25 Sep one is kept in /root/labfix-20260928/replaced/etc/motd; it says “/tmp is cleared at every boot: keep your work in your home directory.”

Back to the table · Requests

H23 On every lab machine, ssh to another lab machine goes over the LAN, not Tailscale#

Fixed 30 Sep updated wastes time all three

2 October Fixed 30 Sep updated Since 30 Sep 15:36 (aifoundry3 corrected)

Fixed on all three: each machine reaches the other two at their tailnet addresses (U24); aifoundry3's wrong line was corrected at 15:36 on 30 Sep.

  • aifoundry1 fixed 28 Sep 20:51: the other two resolve to their tailnet addresses (U24; again at 01:18 on 29 Sep). Before: aifoundry2 and aifoundry3 resolved to LAN addresses (.localdomain): an unpinned ssh aifoundry3 reaches a host with no sshd (H24), and ssh aifoundry2 reaches password OpenSSH over the LAN
  • aifoundry2 fixed 30 Sep 14:37: the other two resolve to their tailnet addresses (U24, run by the owner; checked at 14:41)
  • aifoundry3 fixed 30 Sep 15:36: both peers at their tailnet addresses (the wrong line corrected)

The 25 September write-up below says aifoundry1 and aifoundry3 resolve through MagicDNS; on 27 September they did not. On all three, hosts: in nsswitch is files mdns4_minimal [NOTFOUND=return] dns, and resolved lists the Wi-Fi link's DNS (the router, domain localdomain) next to tailscale0's 100.100.100.100, so the router answers for the bare names (three lookups on each machine, 23:00; again at 23:20).

Next: nothing left

The 25 September write-up, kept as history:

Who acts (25 Sep): The lab (Roman)

What a user sees

From aifoundry2, without a pinned address, ssh aifoundry1 reaches OpenSSH over the Wi-Fi LAN (password logins) instead of Tailscale SSH, and ssh aifoundry3 reaches a machine with no sshd and fails.

What went wrong, and why

known The router's DNS (domain localdomain) answers before MagicDNS on aifoundry2: aifoundry1 and aifoundry3 resolve to LAN addresses. aifoundry1 and aifoundry3 resolve through MagicDNS (corrected on 27 Sep: they do not; see the 27 September state above).

How we found it

The cross-machine consistency check.

What it cost

Confusing failures, and a silent fallback to password SSH over the LAN.

Where it stands

Worked around for our account by ~/.ssh/config pins. The stale hostname line in aifoundry2's /etc/hosts (127.0.1.1 et-Z590-AORUS-MASTER) was fixed by us at 16:24:16 (A14b; now 127.0.1.1 aifoundry2; rollback from /root/labfix-20260925/replaced/).

What the Nekko team can do
  • Add aifoundry1's and aifoundry3's tailnet addresses to aifoundry2's /etc/hosts, or lower the Wi-Fi DNS priority in NetworkManager; tell the other users, since it changes where their ssh goes.
Evidence
  • labfix/consistency.md D11, D12; fix plan C2.10
  • getent hosts on aifoundry2, 17:2x (aifoundry1 at a LAN address, .localdomain)

Re-check, 27 September

  • aifoundry2: getent hosts gave LAN addresses for aifoundry1 and aifoundry3 (.localdomain); our ~/.ssh/config pins the Tailscale addresses
  • aifoundry1 and aifoundry3, 27 Sep 23:00 and 23:20: getent ahostsv4 for the other two machines returns their LAN .localdomain addresses; resolvectl dns shows the router on wlp8s0 next to tailscale0: 100.100.100.100

Back to the table · Requests

H24 aifoundry3 has no sshd: Tailscale is the only way in#

Open blocks work latent aifoundry3

27 September Open Since found 25 Sep

Unchanged: no openssh-server on aifoundry3. Installing it key-only is ours now (U23, waiting for the owner's decision).

Next: Us U23

The 25 September write-up, kept as history:

Who acts (25 Sep): Roman

What a user sees

If tailscaled fails or the tailnet is unreachable, nobody can log in to aifoundry3, root included; the "SSH key fallback over the LAN" in docs/lab-access.md does not exist there.

What went wrong, and why

known openssh-server is not installed on aifoundry3 (it is on the other two).

How we found it

The aifoundry3 audit.

What it cost

Latent: a tailscaled problem means a site visit.

Where it stands

Not changed: it would add a LAN-facing service.

What the Nekko team can do
  • Install openssh-server key-only (PasswordAuthentication no) as a fallback, or document that aifoundry3 needs on-site access when Tailscale fails.
Evidence
  • labfix/audit-aifoundry3.md F22; labfix/consistency.md C6
  • read-only check 17:19 (openssh-server not installed)

Re-check, 27 September

  • aifoundry3: openssh-server not installed; only Tailscale and localhost listeners

Back to the table · Requests

H25 aifoundry2's tmux is a third-party snap that the Ubuntu package would break#

Open workaround wastes time aifoundry2

27 September Open workaround Since 25 Sep 16:24 (hold, us)

Unchanged: the snap tmux stays held. Replacing it after the campaign is ours now (U21).

Next: Us U21

The 25 September write-up, kept as history:

Who acts (25 Sep): Us, then the lab

What a user sees

Installing Ubuntu's tmux 3.4 on aifoundry2 would make tmux attach fail with protocol version mismatch against running sessions of the snap (3.7c), because /usr/bin comes before /snap/bin; a snap refresh can also restart it under running sessions.

What went wrong, and why

known tmux on aifoundry2 is a classic snap from a third-party publisher, not the Ubuntu package.

How we found it

The aifoundry2 audit (our own session runs in it).

What it cost

An innocent install or refresh can kill everyone's tmux sessions.

The workaround

16:24:16 (A14b): snap refresh --hold tmux (held at 17:19); the Ubuntu package is deliberately not installed while our session runs in the snap. Rollback: snap refresh --unhold tmux.

What the Nekko team can do
  • After the measurement campaign, replace the snap with Ubuntu's tmux package, once all sessions have ended.
Evidence
  • labfix/audit-aifoundry2.md U2, R5; labfix/consistency.md C5

Re-check, 27 September

  • snap tmux 3.7c rev 95 held; no snap changes since 25 Sep

Back to the table · Requests

H26 Background load on the measurement hosts, some of it ours#

Open cosmetic all three

4 October Open Since found 25 Sep

It grew again on 2 Oct, and the largest part is ours: our live monitor runs on all three hosts and, every second, checks the holders with sudo and reads each card's temperature through its management node (H35); the dashboard (every 10 minutes) and the History page (every 5 minutes) run from aifoundry2's crontab with probes into the other two hosts, with a watchdog and a node watcher every minute. Pausing all of it during a campaign is ours (U14).

30 Sep: Most of it is ours, and it grew on 30 Sep: a lab dashboard of ours probes all three hosts every 10 minutes from aifoundry2's crontab, runs et-lab-health hourly on each, and reads one card's telemetry per host at most every 30 minutes, behind the holder and lock checks. Our orphaned Chrome is gone (28 Sep, U11); our wireplumber still runs on aifoundry2 and aifoundry3, since the mask of 28 Sep stops only restarts. Pausing the lab's timers and our own probes during a campaign is ours (U21, U14).

Our three orphaned headless Chrome trees on aifoundry2 (24 processes, left by screenshot runs on 27 Sep, with no client connected), whose stop the session's permission check refused on 27 Sep and again at 07:27 on 28 Sep, were stopped at 08:34 on 28 Sep with the owner's approval: a plain kill of the three (20 h and 10 h old, parent PID 1), with their children (U11). At 07:28 we had masked the desktop portal and its GTK backend in our own user manager there (U11), which our screenshot runs kept starting. At 21:20 on 28 Sep we masked the firmware notifier (with its timer, where our user manager had one) and wireplumber in our user managers on all three hosts, which completes U11. Our nodewatch cron and lingering user managers are the owner's call (U14); polling et-holders less often in our queues follows the heat work (U13); pausing the lab's sysstat and plocate timers during a campaign is ours now (U21).

Next: Us H35, U13, U14, U21

The 25 September write-up, kept as history:

Who acts (25 Sep): Us (our cron); the lab (plocate)

What a user sees

tailscaled logs a localapi ping 60 times an hour and cron one line a minute on each machine; aifoundry3 walks the disk daily for plocate; smartd (installed on 25 Sep) reads SMART every 30 minutes.

What went wrong, and why

known Our own nodewatch cron job (every minute) and linger for our account on each machine; plocate on aifoundry3; smartmontools from A4.

How we found it

The aifoundry2 audit's count of journal volume.

What it cost

Small host noise and log volume.

Where it stands

Not changed yet; it is our own load.

What the Nekko team can do
  • We pause nodewatch during rerun windows or record it; the lab may pause plocate-updatedb.timer during measurement campaigns.
Evidence
  • labfix/audit-aifoundry2.md O6; labfix/consistency.md C7, D5; fix plan C1.4

Re-check, 27 September

  • ours: nodewatch cron every minute on all three (4,864 cron lines a day on aifoundry3); lingering user managers; 1,101 of our Tailscale SSH sessions on aifoundry1 since 25 Sep (15,767 tailscaled journal lines on 27 Sep); 4,984 et-holders calls by our queue; an orphaned headless Chrome of ours on aifoundry2 since ~12:30 27 Sep; ~150 xdg-desktop-portal failures a day in our user manager on aifoundry2
  • the lab's: plocate-updatedb and sysstat timers, smartd, the gdm session

Root steps, 28 September (verified 08:00–08:02; the steps after that, by their own output)

  • aifoundry2, 28 Sep 08:34:30: our three orphaned headless Chrome processes (20 h and 10 h old, parent PID 1, 3 children each) stopped with a plain kill, their children with them; at 08:34:44 none of the three was left, and at 08:36:59 and 08:52:37 ps -u lists no Chrome process of ours (labreport2/fix28-aifoundry2.txt)

Back to the table · Requests

H27 An idle login blocks anyone who follows the "nobody else logged in" etiquette#

Open workaround wastes time aifoundry1, aifoundry3

4 October Open workaround Since 18 Sep 17:23

It recurs with the new users: on 4 Oct two newcomers' logins had been idle since 2 Oct (on aifoundry1 and aifoundry3), and two of ours on aifoundry1. Coding agents now run in tmux by design, so only the card lock can be the rule (PO2, H36).

30 Sep: The same idle session on aifoundry1, from 18 Sep, was still there on 30 Sep at 14:22, 12 days on, holding up the reboot (H6); it ended when the host went down at 14:41 (C28). The problem remains until the card lock is the rule (PO2).

Next: Nekko team PO2; us close our idle sessions

The 25 September write-up, kept as history:

Who acts (25 Sep): The lab: etiquette wording

What a user sees

Scripts that follow the etiquette ("check who is logged in before touching a card") never start on aifoundry1: another user's terminal session has been logged in, idle, since 18 Sep. Our smoke runs stopped at 17:13 with others present.

What went wrong, and why

known who cannot tell an idle session from card use, and our own etiquette (CLAUDE.md) was phrased as "check uptime/ps for other users".

How we found it

Our aifoundry1 smoke runs refused to start.

What it cost

Card time lost on aifoundry1 until the rule was changed.

The workaround

By us at 17:23 (amendment A3): on aifoundry1 our queue ignores a user who is only logged in, and blocks on actual card use instead: a node or lock held by another user (et-who), another user's device process, or a CI job.

What the Nekko team can do
  • Make the lab rule "a card is busy when et-who shows a holder or its lock is held", keep the Discord "using / released" convention, and ask users to take /run/lock/etsoc-shire<N>.lock for card work.

Requests: PO2

Evidence

Re-check, 27 September

  • aifoundry1: the same idle session (tmux since 18 Sep 17:23, one long-running process) 9 days on; no other login and no other card user since 25 Sep; it also holds up the H6 reboot

Back to the table · Requests

H28 aifoundry2's card never cools below the governor's 65 °C threshold, so its DVFS is almost never seen#

New corrupts results aifoundry2 (desktop chassis)

4 October New Since 25-26 Sep campaign; 27 Sep (found)

Out of service. The card has been off the PCIe bus since 12:02:01 on 2 Oct (its link reads Unknown, width 63), and the host has stayed on since 10:53 that day with the card in it. Its own temperature cannot be read; the host's drive, network-chip and CPU sensors have tracked aifoundry3's within about 2–3 °C since about 12:30 that day, so there is no sign that the card still heats. There is no workaround: only the cooling fix (SH5), then a power-cutting reboot to bring the card back (C32).

2 Oct: On 1 October at 09:42 this card dropped off the PCIe bus while idle; a full-reset power cycle at 06:45 on 2 October brought it back, and it failed again at 07:42 (126 °C, 103 W). We then watched a whole failure, reading it once a second (full reset at 10:53, host idle, card unused): 45 °C to 85 °C in 22 minutes, a slow crawl from 86 to 94 °C over half an hour, then a runaway: 105 °C at 11:58, 126 °C at 12:01:31, 138 °C (peak sensor reading 144 °C) at 134 W at 12:02:01, and off the bus. The shape is thermal runaway: the idle card's power grows with its temperature (13.6 W plus a leakage part that doubles every 23 °C; 0.44 W rms over 22 to 134 W), and its cooling only just fails to keep up (heat into air at about 51 °C, about 0.9 °C/W). The crawl is a tipping point: at about 97 °C heating and cooling are within 0.1 W of each other, and the case air warming by 3–4 °C over the hour (the host's drive and network-chip sensors) is enough to push the card over; past it nothing can stop it. aifoundry3's identical card, same board and slot, idles at 53 °C and 24 W, which on the same curve means air at about 31 °C. Air 3 °C cooler would hold aifoundry2's card at about 81 °C. The fix is the air reaching this card (SH5). Data: the repository's docs/reports/data/2026-10-02-idle-runaway-aifoundry2/.

Next: Nekko team SH5 (the next visit)

Who acts: On site (airflow); the lab (the per-card sheet)

What a user sees

aifoundry2's card rests at 66–77 °C and runs at 600 MHz, although its governor works. 700–800 MHz appears only after a long idle from cold (once, E10).

What went wrong, and why

known Through the whole version-3 campaign (25–26 Sep) its mean die temperature never read below 65 °C in 336,070 samples; blocks started at 67–99 °C, and 900 s of idle after heating cooled it only to 67–77 °C. The governor steps down above 65 °C and up below it, so a card that never cools below 65 °C stays at its bottom point. inferred The chassis's airflow sets this: aifoundry3's card idles at 55–57 °C.

How we found it

Choosing a card for the heat-placement experiment (27 Sep): the card rested at 67 °C at 16:50 PDT and at 66 °C at the one allowed re-check at 18:52, both already inside its thermal loop, so no clock step could be observed. With aifoundry3 latched (C23) and card 1 off (C24), no card could answer whether the governor trips on the mean or on the hottest sensor; the answer rests on the source alone.

What it cost

The only card whose governor works almost never shows it, so every clock-step experiment needs a cold start after a long idle. Results on aifoundry2 depend on its idle temperature, that is on the room and the chassis (D3).

Where it stands

Worked around by us: blocks preheat to 76 °C so the clock is always 600 MHz (the opposite of what a governor experiment needs), and the heat session allowed one re-check a day.

What the Nekko team can do
  • On site: improve the airflow around aifoundry2's card (chassis fans, or an open bench), then check that it idles below about 60 °C, and log the inlet temperature. It fits the same visit as card 0's check (C21).

Requests: SH5, DI2

Evidence
  • heatplace/feasibility.md l.28, 93–94, 112
  • heatplace/a2-session1.log (16:51, rest 67 °C: WARM) and heatplace/a2-session2.log (18:52, rest 66 °C: WARM)

Back to the table · Requests

H29 aifoundry3's host copies memory at about half the other hosts' rate: it runs on one memory channel#

New workaround corrupts results aifoundry3

27–28 September New workaround Since 27 Sep (E50)

Cause found on 27 Sep at 22:18 and confirmed as root on 28 Sep (dmidecode): one 32 GB DDR4-2666 DIMM, in ChannelA-DIMM1, three slots empty; a second DIMM is on-site work.

Next: Nekko team SH7

Who acts: On site (a second DIMM)

What a user sees

A host memcpy of 256 MB runs at 9.2 GB/s on aifoundry3, against 17.4 GB/s on aifoundry2 and 21.4 GB/s on aifoundry1. A program's staged copy to the card reaches 5.20 GB/s there, against 7.15 and 7.79.

What went wrong, and why

known aifoundry3 has one 32 GB DDR4-2666 DIMM, in ChannelA-DIMM1, and its other three slots are empty, so it runs single-channel memory. aifoundry2 has two DIMMs on two channels (DDR4-2666) and aifoundry1 four on two channels (DDR4-3200). Read on 27 Sep at 22:18 from the memory table that udev exports, as a user.

How we found it

E50 (27 Sep): the pre-registered staged-copy predictions (P4a, P4b) failed on aifoundry3 only, while the DMA-only rates agreed across hosts. The memory layout was read the same evening.

What it cost

Host-to-card throughput, and any host-side timing that moves memory, on aifoundry3 are not comparable with the other hosts.

Where it stands

Cause found on 27 Sep (U10). Compare DMA-only rates, which agree across hosts. The layout is added to the experiment's hosts.txt in our working copy.

What the Nekko team can do
  • Add a matching 32 GB DDR4-2666 DIMM in channel B at the next site visit (request SH7); we then re-run the copy test.

Requests: SH7

Evidence
  • results.md of E50 (P4a, P4b)
  • labreport2/rung30-user-aifoundry3.txt (the DIMM table, 27 Sep 22:18)

Root steps, 28 September (verified 08:00–08:02; the steps after that, by their own output)

  • aifoundry3, 28 Sep 07:26 (dmidecode -t 17, as root): one 32 GB DDR4-2666 DIMM in ChannelA-DIMM1, three slots empty, 64 GB maximum; aifoundry2 07:14: two 32 GB DDR4-2666 DIMMs, one per channel

Back to the table · Requests

H30 et-who prints a sentence when nobody holds a card and always exits 0, so a script took “free” for “held”#

Fixed 28 Sep wastes time all three (our et-who, installed on 25 Sep)

28 September Fixed 28 Sep Since 28 Sep 07:09–07:54 (us, U2)

Fixed on 28 Sep: et-who --check (0 free, 1 held, 2 failed) and the new idle sentence are installed on all three hosts; plain et-who still prints the same holder lines and exits 0.

Next: nothing left

Who acts: Us (fixed on 28 Sep: installed as root at the owner's request)

What a user sees

et-who prints No process has an ET-SoC-1 device node open. when nothing is held, and exits 0 whether or not anyone holds a card.

What went wrong, and why

known Our 25 Sep et-holders (fix A12, C11) echoes that sentence when it finds nothing; it also covers the lock files, which the sentence omits. A caller has to parse text to tell free from held.

How we found it

27 Sep 21:10 PDT on aifoundry3: our heat queue's starter read the idle sentence as a holder and refused to start round 3 ("not starting R3: holders or a queue present"); we started it by hand at 21:12.

What it cost

2 minutes for us. A looser parser could do the opposite and start on a held card.

Where it stands

Worked around: keep only the lines that start with /dev/et or lock: (our lib.sh does). The fix is written and staged on every host (U2): et-who --check exits 0 when free, 1 when held and 2 when the check failed, and the idle sentence becomes “No process holds an ET-SoC-1 device node or card lock.” Plain et-who keeps its output and exit status. Installing it needed root, which the 27 Sep session did not have. Fixed on 28 Sep: installed as root on all three hosts (aifoundry2 07:09, aifoundry3 07:26, aifoundry1 07:54), exactly as tools/lab/README.md says, after a sandbox test showed the old and new holder lines identical. The 08:00–08:02 check confirmed the free, held and failed exit codes, and that the frozen experiment code's parse still finds the holder lines.

What the Nekko team can do

Nothing needed from the Nekko team.

Evidence
  • the queue monitor, 27 Sep 21:10:19 PDT: "not starting R3: holders or a queue present" (session log)
  • labreport2/tools/et-who and labreport2/tools/et-holders (the 27 Sep versions; also in the working copy's tools/lab/)

Root steps, 28 September (verified 08:00–08:02; the steps after that, by their own output)

  • all three, 08:00–08:02: the installed et-who and et-holders match tools/lab/ by sha256; as nobody and as the user, et-who exits 0, --check 0 when free, a bad argument 2
  • aifoundry2 08:01:34 and aifoundry3 08:01:48: with card 0's lock file held for 2 s by flock -n (no card node opened), --check exits 1 and prints 2 lock: lines, plain et-who exits 0 with the same lines, and the frozen experiment code's parse (lines starting with /dev/et or lock:) finds 2; after release, --check exits 0. On aifoundry1 the held test was skipped because another account was logged in
  • before the install, a sandbox test on each host (a private lock directory): the old and new holder lines were identical

Back to the table · Requests

H31 Other accounts now work on aifoundry3, and one used its card; the report assumed only we did#

New workaround corrupts results aifoundry3

4 October New workaround Since 27 Sep 04:35

Several users at once is now normal: three people were onboarded on 2 Oct, and since then two other accounts have used aifoundry3's card and one aifoundry1's card 1 (the card-usage log). The newcomers took the card lock every time; one established account did not (PO2).

27 Sep: Other accounts now log in to aifoundry3: one on 27 Sep 04:35–05:13 (a kernel-launch error followed at 04:42; the card recovered), and a second, different account at 22:43, after our queue ended.

27 Sep, 22:43 PDT: a second account, not the one of 04:35, opened a remote login session on aifoundry3 from the tailnet, after our heat queue R3 had ended; its session was closing at 23:20. No process held the card at 22:58 or at 23:21 (et-who). Our queues on aifoundry3 check et-who and take the card lock before every block.

Next: Nekko team PO2, DI1

Who acts: Roman (make the card lock the rule); every aifoundry3 user

What a user sees

On 27 Sep, 04:35–05:13 PDT, another account was logged in to aifoundry3 over Tailscale SSH, and at 04:42:36 the card logged Minion Runtime Error / OPS API Kernel Launch Error (error code: 2) (MinionCeEvent 1) while none of our work ran.

What went wrong, and why

known aifoundry3's card is shared; the 25 Sep fix plan and our scheduling assumed nobody else used it. Whether that session took the card lock is unknown. The card recovered without a reset: our PCIe runs at 14:49 and the heat runs at 17:00–21:28 passed.

How we found it

The 27 Sep re-check: the tailscaled and kernel journals, and the driver's error counter.

What it cost

No collision with us. An unlocked run by anyone can land inside a measurement, and programs built without the log-level line are exposed to C17 on this host's runtime.

Where it stands

Worked around by us: our queues check et-who and take the card lock. Our aifoundry3 banner says several accounts use the card (U5, installed on 28 Sep), and the onboarding draft (appendix A) covers the lock and C17's line.

What the Nekko team can do
  • Tell every aifoundry3 user about et-who, the card lock and C17's one-line workaround, in the banner and the onboarding page.
  • Make the card lock the rule for everyone (request PO2).

Requests: PO2, DI1

Evidence
  • aifoundry3 journalctl -u tailscaled (login 04:35:36, end 05:13:30), journalctl -k at 04:42:36 and err_stats/ce_count (labreport2/raw-aif3-root.txt)

Back to the table · Requests

H32 aifoundry3's journal has reached its 2 GB cap and will start deleting the July boot history#

Fixed 28 Sep wastes time aifoundry3

4 October Fixed 28 Sep Since 28 Sep 07:11 (us, U9)

Fixed on 28 Sep, and at risk again: since 2 Oct the journal grows by about 140–225 MB a day, about ten times the rate of 28 Sep, mostly sudo lines from our own live monitor (H35). At 2.6 GB of the 4 GB cap it fills around 11–13 Oct and then deletes the July boots again (the export of 28 Sep survives).

28 Sep: Fixed on 28 Sep: the whole journal exported at 06:44 (48 boots back to 17 Jul 01:14, a root-only copy), then our cap raised to 4 GB at 07:11. In between, journald deleted the two oldest files (the 17 Jul 01:14 and 01:18 boots); they survive only in the export.

Next: us H35

Who acts: Us (fixed on 28 Sep: exported, and our cap raised, as root at the owner's request, U9)

What a user sees

journalctl --disk-usage reports 1.8 GB against the 2 GB cap set on 25 Sep (SystemMaxUse=2G); the journal keeps 48 boots, back to 17 Jul.

What went wrong, and why

known Our 25 Sep journal setting (H9) capped aifoundry3's journal at 2 GB, below journald's default of 4 GB for this 457 GB disk, and it has filled. The journal was already persistent: its files go back to 17 Jul. Vacuuming deletes the oldest files first, which hold the 17 and 23 Jul boot clusters that H7 relies on. The desktop noise (H33) speeds it up. The disk is only 22% used. The oldest files are from 17 Jul 01:17, the first of that day's power-cut boots; the active file grows by about 17 MB a day, so the next rotation and vacuum is days away.

How we found it

The 27 Sep re-check.

What it cost

The only on-host record of the July power cuts disappears within days.

Where it stands

Partly saved on 27 Sep at 22:17: as a user we copied aifoundry3's wtmp reboot history, which goes back to 2 Jan (125 boots and only 10 clean shutdown records; the most boots without one on 17 Jul, 13, and 23 Jul, 11). The journal's own record of each July boot needed root to export (U9, staged), which the session did not have. Fixed on 28 Sep, as root at the owner's request (U9): at 06:44 we copied the whole journal (102 files, 48 boots back to 17 Jul 01:14) to a root-only directory, /root/labfix-20260928/journal-export-20260928/, and at 06:50 saved each boot's last kernel and PID-1 lines to our home directory; at 07:11 we raised SystemMaxUse in our 60-labfix.conf to 4 GB and restarted journald. Between the export and the cap, sooner than the estimate above, journald deleted its two oldest files (the 17 Jul 01:14 and 01:18 boots); they survive only in the export. The live journal now keeps 46 boots, back to 17 Jul 01:21, with 1.8 GB of its 4 GB used.

What the Nekko team can do

Nothing needed from the Nekko team: the cap is our own 25 Sep setting, and it is raised.

Evidence
  • aifoundry3, 27 Sep: disk usage, --list-boots, oldest archived files 17, 18 and 21 Jul (labreport2/raw-aif3-root.txt)
  • the wtmp copy, 27 Sep 22:17 (labreport2/FIXES-DONE.json, U9)
  • aifoundry3, 27 Sep 23:20, as the user: SystemMaxUse=2G in 60-labfix.conf; 102 journal files, the oldest system@…journal~ of 17 Jul 01:17; df / 22% used

Root steps, 28 September (verified 08:00–08:02; the steps after that, by their own output)

  • aifoundry3 06:44: /root/labfix-20260928/journal-export-20260928 (mode 0700): 102 files, 2,239,758,336 B, the same count and bytes as the live journal; journalctl -D on the copy lists the same 48 boots from 17 Jul 01:14
  • 07:08: every archived file byte-identical to its copy; one archived file of 23 Jul fails journalctl --verify (“Bad message”) in the live journal and in the copy alike
  • 06:50: each boot's last kernel and PID-1 lines saved to our home directory (100 files; copied to aifoundry2 at 07:27)
  • 07:11:29: SystemMaxUse=4G in our 60-labfix.conf; journald restarted: “System Journal … is 1.8G, max 4.0G, 2.1G free”
  • verification 08:00: the drop-in reads SystemMaxUse=4G and systemd reads it that way; journald active since 07:11:29 with 0 restarts; nothing vacuumed after 07:11; the live journal keeps 46 boots, the oldest from 17 Jul 01:21; the 2 files deleted between 06:44 and 07:11 exist only in the export, and the 96 archived files in both are byte-identical (cmp)

Back to the table · Requests

H33 Desktop services on the headless hosts fill the error log#

Partly fixed workaround cosmetic all three

4 October Partly fixed workaround Since 25 Sep (bluetoothd after the upgrade); earlier for the rest

The desktop part holds on all three: bluetooth, cups-browsed and the updater are off, and all three have booted headless since. The error counts have not been re-read (that needs the adm group). The new noise is ours: our live monitor's core dumps and its per-second sudo lines (H35).

30 Sep: 850–925 bluetoothd errors a day on aifoundry1 until 28 Sep, and the notifier failed every 3 h for each logged-in user on all three. Bluetooth, cups-browsed and the firmware-updater snap are off on all three hosts since 30 Sep (aifoundry1 since 28 Sep; U20), and all three boot headless from their next boot (U27), which removes the greeters. Our own portal, notifier and wireplumber are masked in our user managers (U11). The error counts are to be re-read after a day.

Next: Us H35; re-read the error counts (root or adm)

Who acts: Us, as owner-approved lab fixes since 28 Sep: the headless target (U27) and the desktop services (U20), done on aifoundry1 on 28 Sep at 20:51 and waiting for the owner's decision on the other two; our own user managers (U11, done)

What a user sees

journalctl -p err is mostly noise: bluetoothd logs 850–925 audio-endpoint errors a day on aifoundry1 and about 200 on aifoundry3, the firmware-updater snap's notifier fails every 3 hours for each logged-in user on all three, and on aifoundry2 our own user manager fails the desktop portal about 150 times a day.

What went wrong, and why

known The hosts boot a desktop (H8). The bluetoothd errors started with the 25 Sep upgrade; the notifier fails on an AppArmor denial; cups-browsed added the Wi-Fi LAN's printer on aifoundry3 (localhost only); and every Tailscale SSH login on aifoundry2 logs a PAM error for the missing pam_lastlog.so.

How we found it

Counting error-priority lines per day on each host (27 Sep): 367 of 371 on aifoundry2 since 25 Sep 18:00 are the portal and the notifier.

What it cost

Real errors drown in journalctl -p err, and the noise fills aifoundry3's journal faster (H32).

Where it stands

Partly fixed. Masking the portal and the notifier in our own user managers (U11) was not allowed in the 27 Sep session; only aifoundry2's user manager shows those failures (262 lines in 24 h). On 28 Sep at 07:28 we masked the portal and its GTK backend there, which our screenshot runs' Chrome kept starting. 28 Sep, evening: at 20:51, with the owner's approval, bluetooth, cups-browsed and the firmware-updater snap were switched off on aifoundry1 and its default target set to headless for its next boot, which ends its greeter (U20, U27); at 21:20 we masked the notifier and wireplumber in our user managers on all three hosts (U11). aifoundry2 and aifoundry3 wait for the owner.

What the Nekko team can do
  • Nothing since 28 Sep: the owner has root, so making the machines headless (U27) and turning off bluetooth, cups-browsed and the notifier (U20) are ours: done on aifoundry1 on 28 Sep at 20:51, waiting for the owner's decision on the other two.
  • Leave pam_lastlog alone: the line is optional, and a bad PAM edit breaks every login.
Evidence
  • error-priority lines per day on each host, 27 Sep (labreport2/host-aifoundry2.json and the other two)

Back to the table · Requests

H34 et-who --check covers the whole host and cannot see a card that is down#

New workaround wastes time aifoundry1 (two cards); aifoundry2

4 October New workaround Since 28 Sep (et-who --check); two users on aifoundry1 since 2 Oct

et-who --check exits 1 if anyone holds any card node or lock on the host. On aifoundry1, where two people may now work at once, one per card, a script that gates on it waits for the other person's card. It lists holders only, so a card that fell off the bus keeps its /dev nodes and shows as free (C32). Each call runs sudo (H35). The banners and the onboarding brief say to read only your own card's lines and to check the link in sysfs.

Next: us et-who --card N with a link check (installing it needs root)

Who acts: us (et-who is ours, installed in /usr/local/bin)

What a user sees

et-who --check answers for the whole host: on aifoundry1 it exits 1 when the other card is in use, and on aifoundry2 it exits 0, “free”, for a card that is off the bus.

What went wrong, and why
  • known It was written for one card per host and one user at a time (H30). It lists processes holding a node and lock files held; it does not read the card's link state, and it runs sudo -n et-holders each time.
How we found it

Onboarding newcomers on 2 Oct, one per card on aifoundry1, and aifoundry2's card failing the same day (C32).

What it cost

Waits and wrong “free” answers; once our live monitor called it every second, a sudo line a second in every log (H35).

Where it stands

Worked around in the banners and the onboarding brief since 2 Oct: read only the lines for your own card (/dev/et<N>_ or lock:etsoc-shire<N>.lock), and check the card's current_link_speed. Ours to fix: et-who --card N, a link check that reports “down”, and a holder check that does not need sudo on every call.

What the Nekko team can do
  • Nothing beyond CF5 (a management path that names its holder) and PO2 (the lock as the rule).

Requests: CF5, PO2

Evidence

Back to the table · Requests

H35 Our live monitor reads every card once a second without the card's lock, crashes now and then, and logs a sudo line every second#

New wastes time all three (cards read on aifoundry1 and aifoundry3)

4 October New Since 2 Oct 16:05 (the live collector)

Since 16:05 on 2 Oct our live collector, a user service on all three hosts, reads each card's temperature once a second with a copy of ettelem, but only while et-who --check finds no holder on the host. Each read holds the single-opener management node for about 4 ms without the card's lock: about 177,000 opens of each of aifoundry1's cards by 4 Oct, and a tool started in that window gets “busy”. The reader has crashed in the vendor's libDM.so six times on aifoundry1 and twice on aifoundry3 (no reading lost), the failure C10 warns of. And each tick runs sudo, so aifoundry1 and aifoundry3 log about 3,600 sudo lines an hour, and their journals grow by 150–225 MB a day (H9, H32).

Next: us take the card lock around each read, read the holders without sudo, report the crash (CF5)

Who acts: us

What a user sees

Now and then a management tool started on a free card fails with “Device or resource busy”, about one start in a hundred (the onboarding brief's Traps). The card-usage log shows the opens as unattributed or under someone else's lock. On aifoundry1 and aifoundry3, /var/log/auth.log and the journal fill with one et-holders sudo line a second.

What went wrong, and why
  • known At the owner's request, our live collector (~/live/live-collector.py) reads each card's temperature every second with lab-monitor temp, a copy of ettelem, so the dashboard's Live section and the History page can show it. It reads only while et-who --check exits 0 for the whole host, and it does not take the card's lock.
  • known Each read opens the card's management node for about 4 ms. The card-usage log counted about 177,000 such opens of each of aifoundry1's cards, and as many of aifoundry3's, between 2 Oct and 11:45 on 4 Oct.
  • known The reader crashed with SIGSEGV in libDM.so six times on aifoundry1 and twice on aifoundry3 (coredumpctl), and ettelem itself once on aifoundry1; the records around each crash are normal.
  • known et-who --check runs sudo -n et-holders: 42,253 such sudo lines on aifoundry1 between midnight and late morning on 4 Oct, against 83 on aifoundry2, whose card is down.
How we found it

The 4 Oct audit of this page: the card-usage log, coredumpctl, the journal sizes and auth.log on each host, all read-only.

What it cost

Occasional “busy” failures for other users, core dumps that fill the error log (H33), and journals that now keep about ten days on aifoundry1 (about five at this rate) and will reach aifoundry3's cap around 11–13 Oct (H9, H32). It also contradicts what this page said of our monitor until 4 Oct (CF6, H26).

Where it stands

Ours to fix: wrap each read in the card's lock without waiting (flock -n), and skip the tick when it is held; since a lock-taker using flock -n (et-reset, our burn scripts) could then be refused for those 4 ms, a short wait there is safer. Read the holders without sudo, or less often. Find the libDM crash and report it (CF5). Pause the monitor during a campaign.

What the Nekko team can do
  • A temperature anyone can read without the single-opener node (CF6) would make our monitor unnecessary.

Requests: CF6, CF5

Evidence
  • ~/live/live-collector.py (the same file on all three hosts; LIVE_CARD_EVERY=1), read 4 Oct
  • et-usage --since 2026-10-02 --card 0 on aifoundry1: about 176,800 unattributed opens by 11:44 on 4 Oct; coredumpctl on aifoundry1 and aifoundry3
  • /var/log/auth.log sizes and the journal sizes by day on each host, 4 Oct

Back to the table · Requests

H36 The “is anyone else here” checks miss people, and nothing holds a card between runs#

New workaround wastes time all three

4 October New workaround Since 28 Sep (seen); multi-user since 2 Oct

who reads utmp, which has no entry for a Tailscale SSH command without a terminal: on 28 Sep it listed nobody on aifoundry3 while uptime counted three users. loginctl misses coding agents in tmux under linger: on 30 Sep aifoundry2 showed no sessions while two users had processes running, and every account made since 2 Oct has linger. The card lock lasts one run, so a newcomer setting up or between runs holds nothing.

Next: Nekko team PO2; us our repository docs

Who acts: Roman (the rule); us (our docs)

What a user sees

The old etiquette, “start only when nobody else is logged in”, fails both ways: an idle login blocks everyone (H27), and an active user, or their coding agent, can be invisible to who and loginctl.

What went wrong, and why
  • known who reads utmp, and a Tailscale SSH command without a terminal makes no utmp entry. On aifoundry3, where every login comes through Tailscale, who printed nobody on 28 Sep while uptime counted three users.
  • known loginctl lists sessions, and a coding agent in tmux under a lingering user manager has none: on 30 Sep aifoundry2 showed “No sessions” while two users had processes running. Every new account has had linger since 2 Oct.
  • known The card lock is held for one run, often under 10 s, so it says nothing about someone who is setting up or between runs.
How we found it

Checking who else was on the hosts before our runs, 28 and 30 Sep, and onboarding newcomers on 2 Oct.

What it cost

Wrong answers to “is anyone else here”, in both directions.

Where it stands

Our brief has done this since 2 Oct: count people by the owners of running processes, check the card-usage log for the last 30 minutes, and post the host and card on #community-lab. Our repository's CLAUDE.md, AGENT.md and getting-started.md still say to look with who; correcting them is ours.

What the Nekko team can do
  • Say that et-who and the card lock decide, not logins, and give a way to claim a card for a working session with a time limit, for example a reservation that et-who shows (PO2).

Requests: PO2

Evidence

Back to the table · Requests

H37 The root installs staged on 2 October were never made: two login banners contradict the cards' state#

New wastes time all three

4 October New Since 2 Oct 14:24 (staged, not installed)

aifoundry1's login banner (30 Sep) still says card 0 overheats and must not be used, though its fan was replaced on 2 Oct and the dashboard and the New user page offer it. aifoundry2's does not say its card is out of service, and still invites newcomers to et-lab-start and the card lock. The installed et-lab-start (2 Oct 14:33, all three) hands the agent over at step 4 of the brief, which skips the simulator check. aifoundry1's card-usage logger skipped card 0 until 4 Oct; it logs it again. Corrected copies have been staged since 2 Oct and need one root session (U28).

Next: us U28 (root)

Who acts: us, with root (the banners and et-lab-start are ours, U5)

What a user sees

A newcomer logging in to aifoundry1 reads “card 0 OVERHEATS … Do not run work on card 0” while the New user page sends them to card 0; on aifoundry2 the banner offers et-lab-start and the card lock for a card that has been off the bus since 2 Oct (C32, H28).

What went wrong, and why
  • known On 2 Oct we corrected the banners (tools/lab/motd-aifoundry1 and motd-aifoundry2), et-lab-start, a notice for root logins and the et-opens tracer in the repository, and staged them on the hosts; each needs root to install, and the root session did not happen.
  • known The installed et-lab-start (2 Oct 14:33) tells the agent to follow the brief “from step 4”, which skips its simulator check; the repository's copy fixes that line.
  • known aifoundry1's card-usage logger still runs with the options that excluded card 0 while it was out of use, so card 0's queue activity is not logged.
How we found it

The 4 Oct audit: the installed /etc/motd and /usr/local/bin/et-lab-start on each host against the repository's and the staged copies (checksums).

What it cost

Contradicting instructions for the people onboarded on 2 Oct.

Where it stands

Staged and checked; waits for one root session (U28). C21 had named the Nekko team for aifoundry1's banner; the banners are ours (U5).

What the Nekko team can do
  • Nothing: this is ours.
Evidence
  • aifoundry1 /etc/motd (30 Sep 22:37) and aifoundry2 /etc/motd (30 Sep 20:55), read 4 Oct; the staged copies equal the repository's (tools/lab/)
  • /usr/local/bin/et-lab-start on all three hosts, 2 Oct 14:33, older than the repository's

Back to the table · Requests

3.3 Developing and measuring#

What trips up someone writing kernels for the chip or measuring power and time on it, and the day-to-day mechanics of working on the hosts. 25 problems on 4 October: 3 fixed, 1 partly fixed, 16 open, 5 new; 2 of the open and partly fixed ones updated since 27 September.

D1 The card's meters are coarse and filtered, and half of an idle card's power is on no meter#

Open updated workaround corrupts results all four cards

30 September Open updated workaround Since 26 Sep (E41)

Numbers replaced by E41, and the filter by E58 (29 Sep, pre-registered): a first-order average of 1.01–1.06 s on aifoundry3's rails and 1.08 s on card 1's minion and NoC rails, but 0.54 s on card 1's SRAM rail, so the cards' meters differ; aifoundry3's slower pass has a candidate cause (C23).

E41 (26 Sep, three cards, three passes each) replaced the single-session numbers: with nothing polling the SP pass is 133.2 ms (aifoundry2), 134.8 (card 1) and 224.1–224.5 (aifoundry3); under ettelem at 10 Hz 156, 158 and 263 ms; at 25 ms 307–317 and 582 ms — sampling harder returns fewer fresh readings. The 'why is aifoundry3 slower' unknown now has a candidate (C23). The stock firmware has no per-shire temperature readout at all (the per-shire print has no caller), so the 27 Sep heat-placement experiment could only time the 34-sensor mean; a stats reset restarts the rail averages (C12).

Next: Nekko team CF11, DI3

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko: firmware telemetry and documentation

What a user sees

Short bursts read low. Rail powers lag a step by about 1.2 s; the temperature reads whole degrees and is the mean of 34 sensors; board power refreshes only every ~133 ms on aifoundry2 and ~250 ms on aifoundry3, however fast you poll; one sample (six management commands) takes about 22 ms, so 45 Hz is the ceiling. Board power minus the three metered rails is about 15 W at idle (of 31.8 W at 73 °C) and about 21 W under load, and DRAM traffic barely shows on the rails. The stock CLI does not expose the per-rail snapshot (DM_CMD_GET_SP_STATS).

What went wrong, and why

known (measured). The service processor forwards the PMIC's running average of each output rail: roughly first order with tau 1.15–1.22 s, i.e. 55–57% of a step after 1 s, 83–84% after 2 s and about 94% after 3 s. Only the minion, SRAM and NoC regulators are forwarded; DDR core, VDDQ, VDDQLP, PCIe logic, PCIe/PShire, the IO shire and the Maxions have set points and no telemetry. The 35 individual temperature sensors, the PMIC's input-side readings and the process detectors are not exported. unknown Why aifoundry3 refreshes half as often as aifoundry2 (inferred on the card: its service-processor loop or its PMIC).

How we found it

Step-response fits over 242 and 229 bursts (E27) and the attribution of the unmetered remainder over about 390 configurations on both cards (E30), 20–24 Sep.

What it cost

Days of method work before small power differences could be trusted. A new user who averages a 1 s burst gets about half the rail power, and rails alone miss 3/5 to 3/4 of what DRAM traffic adds.

The workaround

By us: bursts of 3 s or more, 2–3 s skipped after each change, a 1/0.94 rail scale, our own sampler (tools/ettelem), sub-degree temperature recovered from the times the reading steps, and cards compared on switching power over idle. A fit attributes the unmetered remainder (18–20% of minion-rail power, 68–73 pJ per DRAM byte), and the memory shires' voltage monitor (die_mv.ddr) serves as a free DRAM-activity meter.

What the Nekko team can do
  • Document the filter, each card's refresh period, and which rails the host can see, so nobody takes rail power for card power.
  • Ship a supported telemetry sampler in /opt/et that reads the per-rail snapshot at 10–45 Hz.
  • In firmware (C9): forward instantaneous rail power with a timestamp per reading, the PMIC's input-side power per regulator, and the 35 live sensors. In lab hardware, a PCIe riser with shunts read at a kilohertz would sharpen every small-signal measurement.

Requests: CF11, DI3

Evidence

Re-check, 27 September

  • E41: SP pass with nothing polling 133.2 ms (aifoundry2), 134.8 (card 1), 224.1-224.5 (aifoundry3); under ettelem at 10 Hz 156/158/263 ms; at 25 ms 307-317 and 582 ms: polling harder returns fewer fresh readings
  • the stock firmware has no per-shire temperature readout at all (the print has no caller); a stats reset restarts the rail averages (C12)

Back to the table · Requests

D2 Some on-chip traffic starves the card's own meter on aifoundry2#

Open workaround corrupts results aifoundry2 (aifoundry3 unaffected; aifoundry1 not yet characterised)

27 September Open workaround Since 22 Sep

Unchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.

Next: Nekko team CF11

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko: find the cause

What a user sees

On aifoundry2 a telemetry read (six management commands, normally 22 ms) takes 76–146 ms during rings between shires s and s+16, 0.8–1.6 s during tensor loads between shires three hops apart in one column, and a median of 23–206 ms (683 ms at most) during DRAM traffic. The board value then changes only 1.4–2.4 times a second, and a burst's energy reads about 40% low (33–45%). aifoundry3 ran the same patterns at 21–22 ms.

What went wrong, and why

Effect known, mechanism unknown. inferred The workload's mesh or memory traffic contends with the service processor's path to the PMIC or to the host queue. Why aifoundry3, with the same firmware 1.3.1 and the same 400 MHz NoC clock, is immune is not established.

How we found it

E29 (rings) and E31 (heat per mm), 23–24 Sep: energy readings far below the other passes, traced to the sampler's own latency.

What it cost

Burst energies about 40% low if not caught; six E31 bursts dropped; the rings row of the energy manual reported from aifoundry3 only.

The workaround

By us: every sample records its own read time (took_ms), and the analysis drops bursts whose median exceeds 60 ms. The cause has not been investigated on the card.

What the Nekko team can do
  • Find which traffic delays the service processor's management loop, whether its path can be prioritised, and whether other cards share the problem.
  • Timestamp each reading in firmware so a stale value can be recognised; document each card's refresh period.

Requests: CF11

Evidence

Re-check, 27 September

  • card-side; card, firmware and runtime unchanged since 25 Sep

Back to the table · Requests

D3 Power follows die temperature, heat carries over between runs, and the room's airflow changes#

Open updated workaround corrupts results all (worst on aifoundry2's desktop chassis)

27 September Open updated workaround Since 27 Sep

aifoundry2's chassis keeps its card hot enough to hide its governor (H28).

aifoundry2's chassis keeps its die above the governor threshold: never below 65 °C in 336,070 campaign samples, blocks starting at 67–99 °C, and the 27 Sep heat session WARM at both allowed attempts (67 °C at 16:50, 66 °C at 18:52). See H28 for the site fix.

30 Sep: DV2's validation on aifoundry2 (28–29 Sep, E51) confirmed it with a cost: the idle die read 71–76 °C in 327 of 378 cycles, so two of its heating tests and the placement question stayed untested; the one usable window was an evening cool spell, which fits a change in the room (inferred). Room temperature is still not logged (SH5).

4 Oct: the hosts' CPU, drive and network-chip temperatures are logged every 5 s since 2 Oct (the History page); the room's and the case inlet's still are not.

Next: Nekko team SH5

The 25 September write-up, kept as history:

Who acts (25 Sep): Users (method); the lab (log the room)

What a user sees

The same kernel's energy depends on what ran before it and on the time of day; a 2 W signal reads 10–50% high; board power creeps up about 3 W over 12 s of a steady load. Once a card's die fell from 84 to 72 °C in four minutes under an unchanged workload.

What went wrong, and why

known (measured). Leakage dominates: idle power is 12.6 W + 23.3 W × exp((T − 80)/36), a slope of 0.65 W per °C at 80 °C (0.81 W per °C of drift inside hot 7 s runs on aifoundry2). The leakage loop gain is 0.95 at 80 °C and passes one at 82 °C, so from 80 °C random fp32 reaches 90 °C in about 20 s. inferred The sudden cooling was someone changing the airflow in the lab; the cards sit in desktop chassis and idle at 62–80 °C.

How we found it

Repeating runs at different start temperatures; the idle law fitted from 64 to 88 °C; a thermal fit that broke mid-session (E12).

What it cost

run_energy.py's polling, which recorded no die temperature, read 10–50% high on 2 W signals on a cooling card; the 18 Sep ring and level energies (2–20% high) had to be re-sampled; the protocol needs 3–8× more wall-clock time than card time; the E12 thermal fits had to stop at the airflow change.

The workaround

By us: every burst is bracketed by 4.5–6 s of idle; energy is corrected for leakage with the idle law and the sampler's die temperature; launches start at a fixed falling edge (80.96 ± 0.11 °C) with configurations shuffled into blocks; long runs are capped at 90 °C. The lab side is unchanged (lm-sensors, installed in A4, now gives host temperatures).

What the Nekko team can do
  • Publish each card's idle law and thermal resistance, and say in the lab notes that in these chassis sustained switching power above about 3 W has no thermal equilibrium.
  • Log ambient or inlet temperature next to the machines (a cheap sensor is enough), keep the chassis airflow fixed, and note in a lab log when doors, fans or air conditioning change.

Requests: SH5

Evidence

Re-check, 27 September

  • aifoundry2: mean die never below 65 C in 336,070 samples; blocks started at 67-99 C; WARM at both allowed checks 27 Sep (67 C at 16:50, 66 C at 18:52) -> H28
  • room temperature still not logged; the OS cannot see card temperature or chassis fans (C25)

Back to the table · Requests

D4 Common instructions trap in user mode: divide, square root, sine, 64-bit integer-to-float, double, the cycle CSR#

Open workaround wastes time all four cards (the chip)

27 September Open workaround Since 18-22 Sep

Unchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.

Next: Nekko team RT4, CF11, DI3

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko: toolchain, firmware, errata

What a user sees

The kernel compiles and assembles without complaint, then faults on the card: a float divide, sqrtf(), a double constant, a 64-bit loop counter converted to float, or reading the cycle CSR.

What went wrong, and why

That they trap is known; the mechanism is inferred: no hardware fdiv or fsqrt, and no user-mode emulation in the firmware. Thirteen instructions trapped in a one-off check: fdiv.s, fsqrt.s, fdiv.ps, fsqrt.ps, frsq.ps, fsin.ps, fdiv.pi, fdivu.pi, frem.pi, fremu.pi, fcvt.l.s, fcvt.s.l and csrr cycle; fcvt.s.lu traps with cause 30 (0x1e). A double anywhere in a kernel brings in 64-bit conversions. The toolchain accepts all of them; gp-sdk builds with -mno-fdiv and ships replacements, but a standalone build has to arrange this itself.

How we found it

Writing the instruction catalogue (E26, E27), traceprof and memprobe.

What it cost

Debugging time per trap. The 13-instruction list rests on a one-off check whose card log was not kept, and reproducing a fault on a shared card is itself a risk.

The workaround

By us: 32-bit counters in floating-point loops, no double, -mno-fdiv and replacement routines, and hpmcounter3 instead of the cycle CSR (with the fix in D13).

What the Nekko team can do
  • List the user-mode-trapping instructions in the PRM or the errata, and give the toolchain a CPU option that rejects them at compile time.
  • Have the firmware's trap handler report the cause, PC and instruction to the host ("illegal instruction fsqrt.ps at PC …") instead of a failed launch.

Requests: RT4, CF11, DI3

Evidence

Re-check, 27 September

  • card-side; unchanged

Back to the table · Requests

D5 Scratchpad addressing traps: offset 0 faulted once, and a global atomic through self ID 0x7F is a bus error#

Open workaround wastes time all four cards (the chip)

27 September Open workaround Since 19-22 Sep

Unchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.

Next: Nekko team DI3, CF11

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko: documentation

What a user sees

A kernel faults or raises a bus error on a scratchpad address that looks legal.

What went wrong, and why

(a) Effect known, mechanism inferred: an amoaddg to a scratchpad word addressed through the self ID 0x7F raises a bus error (E22), presumably because the local path does not carry global atomics; the same word addressed by the shire's explicit ID works. (b) Effect known, cause unknown: offset 0 of a shire's scratchpad faulted in E24. Its scope is unclear: memprobe keeps its op list at offset 0 of its own scratchpad through 0x7F, and memhier reads from offset 0, both without faulting, so the rule as our notes state it ("offset 0 faults") may be broader than the evidence.

How we found it

The on-chip relay (E24) and hot-line (E22) experiments, 22 Sep.

What it cost

Both found the hard way. Every tool since starts its buffers 256 KB in, which gives up 10% of each 2.5 MB scratchpad.

The workaround

By us: scratchpad buffers start at 256 KB, and global atomics go to DRAM lines or through the explicit shire ID. The vendor has not been asked.

What the Nekko team can do
  • Document which scratchpad regions are reserved and why (PRM chapter 15), and which address formats support atomics on the L2 scratchpad.
  • Report the faulting address and cause to the host.

Requests: DI3, CF11

Evidence

Re-check, 27 September

  • card-side; unchanged

Back to the table · Requests

D6 The L1 data cache is not coherent and writes back whole 64 B lines#

Open workaround corrupts results all four cards (the chip)

27 September Open workaround Since 19 Sep

Unchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.

Next: Nekko team RT3

The 25 September write-up, kept as history:

Who acts (25 Sep): Users (method); Nekko (library defaults)

What a user sees

Wrong results when harts on different minions write neighbouring data, for example when a matmul's inner loop is parallelised. A TensorStore read by another minion's TensorLoad can return stale data unless the shire's coalescing buffers are drained first.

What went wrong, and why

known By design: up to about 2,465 caches with no coherence between them. A dirty L1 line is written back whole, so two minions writing the same line overwrite each other.

How we found it

The documentation (PRM, the FOSDEM talk) and sys_emu's checkers.

What it cost

Silent wrong answers for anyone porting GPU-style code. sys_emu catches most cases with -mem_check, which the runtime enables by default.

The workaround

By us: each hart owns whole 64 B-aligned lines of its output; L/G memory operations or explicit evicts where data is shared; validation in sys_emu with -mem_check. The firmware evicts L1 and L2 after each kernel.

What the Nekko team can do
  • Make gp-sdk and et-common-libs allocate per-hart outputs line-aligned by default, and put this rule first in the onboarding notes; keep -mem_check on by default.

Requests: RT3

Evidence

Re-check, 27 September

  • card-side; unchanged

Back to the table · Requests

D7 One hot line stops a shire: hammering a global atomic stalls its home shire's memory path, with no error#

Open workaround blocks work all four cards (measured on aifoundry2 and aifoundry3)

27 September Open workaround Since 22 Sep

Unchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.

Next: Nekko team RT3, DI3

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko: library and programming guide

What a user sees

The 32 minions of one shire sit in loads that do not return while other shires hammer a global atomic homed in that shire. A Discord report read this as "shire 0 got 6% of its fair share". A barrier that spins on a global atomic can stall a whole shire.

What went wrong, and why

Effect known; the mechanism is as described in errata 4.1 (RTLMIN-6207) and 4.2 (RTLMIN-6214): at the shire cache, requests from the mesh rank above the shire's own neighbourhood requests. The home shire's scratchpad or DRAM throughput falls to 0.01–0.05% of normal (384 operations while 6 M atomics retire). It is a stop, not a slow-down, and the threshold is a cliff that one other shire's minions clear. The atomic itself is shared fairly (0.998–1.004 of an even split).

How we found it

E22 and E23, re-measuring the Discord report on both cards.

What it cost

A misdiagnosed report on Discord; any GPU-style spin barrier or counter is exposed.

The workaround

By us: nocbench's credit-release barrier, remote atomics paced (the errata's workaround), contended counters spread over 32 lines. Whether gp-sdk's or et-common-libs' own barriers spin on a global atomic has not been checked.

What the Nekko team can do
  • Document this with the two errata in the programming guide; ship a barrier in et-common-libs that does not spin on a global atomic, and audit the existing sync helpers.

Requests: RT3, DI3

Evidence

Re-check, 27 September

  • card-side; unchanged

Back to the table · Requests

D8 VPU register-file erratum 1.29: the compiler inserts no workaround and the simulator does not model it#

Open corrupts results all four cards (stepping unconfirmed)

27 September Open Since 20 Sep

Unchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.

30 Sep: a new case on 29 Sep (E59): the compiler's own save and restore of the callee-saved f registers, around a function whose asm clobbers them, produced a type-A hazard, which sys_emu -vpurf_warn caught before any card run (RT4).

Next: Nekko team RT4, DI3, CD1

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko: stepping, toolchain

What a user sees

Potentially, vector code that is correct in sys_emu gives wrong results only on silicon.

What went wrong, and why

The erratum is known; its impact on the lab cards is unknown. A VPU register read in the cycle after a write can return stale data. GCC 15.1 (the lab's) and 15.2 (upstream) do not insert the workarounds. sys_emu -vpurf_warn flags gp-sdk's own saxpy_vector inner loop and some service-processor and master-shire firmware routines; -vpurf_check aborts on the firmware's own code. On aifoundry3, workloads/sgemm, which has the flagged flw→fmadd pattern on nearly every FMA, gave correct results for all 1.6 M outputs checked, so the hazard is at least not easy to hit.

How we found it

sys_emu's checker on the gp-sdk examples (18 Sep).

What it cost

An unresolved correctness doubt over every vector kernel; kernels have to be checked with -vpurf_warn and the results filtered by PC.

Where it stands

Two questions are unanswered: whether the lab cards are A0 silicon (they report ASIC revision 597), and whether the compiler is meant to handle the erratum.

What the Nekko team can do
  • State each lab card's silicon stepping and whether erratum 1.29 applies.
  • Add the workaround in a compiler pass, or at least an assembler warning; make -vpurf_check skip firmware PCs so it can run by default.
  • Say which of the open RTL's two vector register files the silicon has, so that the erratum can be simulated in the right one (CD1).

Requests: RT4, DI3, CD1

Evidence

Re-check, 27 September

  • card-side; unchanged

Back to the table · Requests

D9 sys_emu and silicon disagree in ways a new user does not expect#

Open workaround wastes time sys_emu (laptop VM and lab hosts)

27 September Open workaround Since 18 Sep

Unchanged: sys_emu and the cards are as on 25 Sep.

30 Sep: sys_emu's VPURF checker ignores tensor writes to the f registers (erratum 1.29 type F), so a clean -vpurf_warn run does not cover that type (E59, RT6).

Next: Nekko team RT6

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko: gp-sdk and sys_emu

What a user sees

Code that passes on sys_emu hangs or misbehaves on a card. A gp-sdk launcher on sys_emu spins forever on Trapping to the same address and writes a sysemu.log0 of several GB. sys_emu gives no usable timing.

What went wrong, and why

known sys_emu is functional only. It models neither erratum 1.29 (D8) nor the single TensorSend ready bit (C14). gp-sdk's GenericLauncher never gives sys_emu the firmware ELFs to preload, so the service processor boots on empty memory. Simulated boot takes about 37 s, and DRAM resets to 0xDEADBEEF.

How we found it

The first hello-world and gp-sdk runs, 18 Sep.

What it cost

Time in the first session, and false confidence about silicon from a simulator pass.

The workaround

By us: patches/et-platform-0002-gpsdk-sysemu-preload-firmware.patch (applied by scripts/provision-vm.sh), the same fix inside patches/lab-gp-sdk-06605ab.patch for the lab hosts, and small shire masks. Whether the preload fix is upstream is not known.

What the Nekko team can do
  • Upstream the firmware-preload fix; publish a list of what sys_emu does not model, next to its checkers; add a checker for TensorSend partner conflicts.

Requests: RT6

Evidence

Re-check, 27 September

  • unchanged

Back to the table · Requests

D10 Current gp-sdk does not build against the lab's /opt/et, and there is no shared install#

Open workaround blocks work all three hosts

27 September Open workaround Since 18 Sep

Unchanged: still no /opt/gp-sdk on any host; it waits for one /opt/et (C18).

Next: Nekko team RT2, RT6

The 25 September write-up, kept as history:

Who acts (25 Sep): Lab: a shared install. Nekko: matching versions

What a user sees

Current gp-sdk requires Erbium components that the lab's /opt/et lacks. The last gp-sdk before Erbium (06605ab), built as it is, produces kernels that all fault at PC 0x40 with an instruction access fault, and its simulator runs have no firmware (D9). Each user builds a private copy.

What went wrong, and why

known The lab's /opt/et is et-platform 353f20e (December 2025; runtime 0.19.0, GCC 15.1), older than the gp-sdk that upstream documents. gp-sdk 06605ab links kernels without -Wl,--emit-relocs, so the runtime cannot relocate the entry-point table (upstream restored the flag in 7e3b4c2). It also needs dnnLibrary and FFTW, and its riscv_helpers.cmake uses a target_link_libraries form that clashes with a user's kernels/CMakeLists.txt. No host has a system-wide gp-sdk; aifoundry3 has several unbuilt source copies whose README builds inside a vendor Docker image.

How we found it

Building our first benchmark on aifoundry2, 18 Sep.

What it cost

A day of patching in the first session; every new user repeats it, and results depend on which commit and patch each used.

The workaround

By us: scripts/deploy-lab-gpsdk.sh exports gp-sdk 06605ab, applies patches/lab-gp-sdk-06605ab.patch and builds under the user's home. A4 added libfftw3-dev on aifoundry1 and aifoundry3. A shared /opt/gp-sdk (fix plan C1.3) was not built: it is absent on all three hosts (17:17), and is best done after the runtime is unified (C18).

What the Nekko team can do
  • Choose a supported gp-sdk and install it in /opt/gp-sdk on every host, built against the same runtime, and name it in the banner. Better, move /opt/et to a release that current gp-sdk supports, and publish which /opt/et and gp-sdk versions go together.
  • Upstream the --emit-relocs and firmware-preload fixes.

Requests: RT2, RT6

Evidence

Re-check, 27 September

  • no /opt/gp-sdk on any host

Back to the table · Requests

D11 Kernel ELFs built on different hosts hash differently, but the code is identical#

Open workaround cosmetic all three hosts

27 September Open workaround Since 22 Sep

Unchanged (cosmetic): the toolchain in /opt/et is as on 25 Sep.

Next: Nekko team RT2

The 25 September write-up, kept as history:

Who acts (25 Sep): Lab (optional)

What a user sees

The same sources give a kernel ELF with a different md5 on each host (sparsity.elf f465… on aifoundry2, 2b56… on aifoundry3). We first recorded this as "the hosts' toolchains build different kernels".

What went wrong, and why

known Each host has its own build of gcc 15.1.0 and binutils 2.45 from the same sources, and they generate the same code: .text, .rodata and .sdata are byte-identical. Only the .comment string differs (GCC: () 15.1.0 against GCC: (g1b306039a) 15.1.0). One real side effect: aifoundry1's lto1 was built without zstd, so it cannot read LTO objects compiled on aifoundry2 or aifoundry3. The lab's compiler (15.1) also differs from upstream's 15.2.

How we found it

Section hashes and cross-compiling the same sources on each host (the consistency check, 25 Sep).

What it cost

A false lesson, since corrected, and time spent comparing ELFs.

The workaround

Compare hashes of .text or of the extracted binary, not of the whole file. Nothing changed on the machines; our repo notes are still to be corrected (D18).

What the Nekko team can do
  • Optionally install one toolchain build on all three hosts (byte-identical ELFs, LTO compatibility), and publish the toolchain commit with /opt/et.

Requests: RT2

Evidence
  • labfix/consistency.md bottom line 1, A14; labfix/audit-aifoundry2.md E7
  • labfix/tc-aifoundry1.txt … labfix/tc-aifoundry3.txt

Re-check, 27 September

  • toolchain in /opt/et unchanged

Back to the table · Requests

D12 A stray write from a kernel leaves no trace on the card or the host#

Open workaround corrupts results all four cards (seen on aifoundry2 and aifoundry3)

27 September Open workaround Since 22 Sep

Unchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.

Next: Nekko team CF11

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko: firmware protection

What a user sees

On 24 Sep a kernel's tensor stores from shire 0 went to physical address 0 (the start of the PU region's Maxion window) in two test launches of about half a second each. No error counter moved, no telemetry field changed, the clock stayed at 600 MHz, and the card ran a normal burst straight after.

What went wrong, and why

The trigger is known (our bug): a new memory mode missing from our host's list launched with no buffer, so the slice base was 0. That nothing noticed is inferred: user-mode kernels are not kept out of that window, DRAM ECC is compiled off, and the SRAM ECC interrupt sources are never enabled, so corruption elsewhere would be just as silent.

How we found it

The kernel's author reread the code afterwards; nothing on the card or host reported it.

What it cost

About a second of writes into a region whose effect on the Maxion side is unknown; the risk is silent corruption of card state or of another user's run.

The workaround

Guarded in our tool only: enercat_host refuses to launch a memory pattern that has no buffer slice.

What the Nekko team can do
  • Keep user-mode kernels out of the PU and Maxion windows (PMP or region checks in the minion firmware), or at least count such accesses in an error counter the host can read.
  • Enable the ECC reporting sources, and document the physical address map for kernel writers.

Requests: CF11

Evidence

Re-check, 27 September

  • card-side; unchanged

Back to the table · Requests

D13 Cycle counting traps: hpmcounter3 reads 128 short, the cycle CSR traps, and evict_va is asynchronous#

Open workaround corrupts results all four cards (the chip)

27 September Open workaround Since 20 Sep

Unchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep.

Next: Nekko team RT3, DI3

The 25 September write-up, kept as history:

Who acts (25 Sep): Nekko: errata and library

What a user sees

Timed loops show occasional −128-cycle outliers; reading the standard cycle CSR faults (D4); a timing taken right after an evict measures the wrong cache level.

What went wrong, and why

known, and the counter bug was reproduced in the original RTL: the PMU's 12 counters share one adder, which folds 7-bit pre-counter overflows into the post-counters round-robin, and a read ignores a pending overflow, so hpmcounter3 reads 128 short whenever its low 7 bits are 0–10. The firmware's four-reads workaround does not fix it. Concurrent PMU reads by both harts can also be wrong (erratum 1.23). evict_va returns before the line has moved; its level code names where the line is left (1 = L2, 2 = L3, 3 = memory).

How we found it

Latency histograms in memprobe, then RTL simulation under Verilator (20 Sep).

What it cost

Wrong latencies until the carry bug was found.

The workaround

By us: fixcyc() adds 128 back; a fence and a wait of a few hundred cycles after evicts; one hart reads the PMU. A firmware syscall that lets kernels choose counter events (patches/0003) works in sys_emu but needs a signed image (C9).

What the Nekko team can do
  • Add the late carry to the errata; ship a corrected counter-read helper in et-common-libs; document evict_va's asynchrony and level codes.

Requests: RT3, DI3

Evidence

Re-check, 27 September

  • card-side; unchanged

Back to the table · Requests

D14 Code on the lab hosts goes stale, or breaks, when it is updated in place#

Open workaround wastes time all three hosts

27 September Open workaround Since 22 Sep

One more instance, fixed and written down; the lab part is three lines in the onboarding page.

One more instance: at 03:17 PDT 26 Sep the gather/scatter queue on aifoundry2 failed to start ('nohup: failed to run command tools/claims-v3/queue.sh: Permission denied'): the script had lost its executable bit. Fixed in the repository (5b968a8). New lessons in AGENT.md §7: never rebuild into a build directory a running queue's blocks use, and do not run the deploy scripts against a host whose queue runs (they rebuild build/ directories).

Next: Nekko team DI1

The 25 September write-up, kept as history:

Who acts (25 Sep): Users

What a user sees

After syncing new sources, make on the host says there is nothing to do, and the old binary runs. A runner script edited or copied over while it runs does something random or aborts. A script copied to a new path is no longer executable.

What went wrong, and why

known rsync -a and tar keep the source files' modification times, so a build directory with newer objects looks up to date. Bash reads a script by offset while it runs, so replacing its bytes in place corrupts the running process. scp without -p drops the executable bit on a new path.

How we found it

The failures themselves, 20–25 Sep.

What it cost

A four-hour session aborted at its last step; aifoundry3's hot-line pass 3 was cut short when its driver script was overwritten; stale-binary runs had to be caught by hashing binaries.

The workaround

By us: builds after a sync use --clean-first; scripts are replaced by writing a temporary file and renaming it (mv), never edited in place; chmod +x after copying; every block records code.sha256 and binary.sha256.

What the Nekko team can do
  • Three lines in the onboarding page or banner: build clean after syncing, rename rather than overwrite a running script, record binary hashes.

Requests: DI1

Evidence

Re-check, 27 September

  • 26 Sep 03:17: the gather/scatter queue on aifoundry2 failed to start because queue.sh had lost its executable bit (fixed in the repository, 5b968a8); new AGENT.md §7 lessons (never rebuild into a build directory a running queue uses; no deploy scripts against a host whose queue runs)

Back to the table · Requests

D15 Long or remote jobs over Tailscale SSH die, kill themselves, or expose their command lines#

Open workaround wastes time all three hosts

30 September Open workaround Since 18 Sep

Hit again on 27 Sep; the trap in our checklist was fixed that night (D22). Command lines are still readable by every user on all three hosts (30 Sep).

Hit again 27 Sep 21:11 PDT: pgrep -f run_queue.sh inside an ssh command matched its own tailscaled be-child process twice while we restarted the heat round on aifoundry3, and AGENT.md §11 then recommended pgrep -af queue.sh over ssh (D22, fixed that night). No Tailscale check prompt was hit on 26–27 Sep. (Our sessions from aifoundry2 go through Tailscale SSH, not the LAN OpenSSH path this note first said: the remote shell's parent is tailscaled be-child, checked on 30 Sep.)

Next: Nekko team AS2, DI1

The 25 September write-up, kept as history:

Who acts (25 Sep): Users; Nekko (onboarding page)

What a user sees

A job started with plain nohup … & over ssh dies when the ssh or its caller times out. pkill -f <pattern> over ssh kills your own session, and pgrep -f <script> always finds a match. Every user on a machine can read the full command line of every ssh command, root's included, and tailscaled writes each one to the journal (about 20,000 lines a day on aifoundry3).

What went wrong, and why

known Plain nohup keeps the ssh channel open. tailscaled runs each command as tailscaled be-child ssh … --cmd=<the whole command>, so the pattern appears in a process that pkill -f matches, and process command lines are world-readable by default. The journal copy is readable by the adm group.

How we found it

We killed our own sessions in the first lab session; the command lines showed up in the fix transcript.

What it cost

Lost runs and lost sessions; command lines visible to everyone.

The workaround

By us: setsid nohup <cmd> < /dev/null > log 2>&1 &, polled with pgrep patterns the ssh command does not contain (tools/claims-v3/queue.sh); processes killed by PID; no secrets on ssh command lines; users deliberately not added to adm (H15). tmux is now on all three hosts (A4 added it to aifoundry3, the one without it).

What the Nekko team can do
  • A short "running long jobs" note in the onboarding page and banner: setsid nohup or tmux, kill by PID, never tokens or passwords on an ssh command line. Keep adm for admins.

Requests: AS2, DI1

Evidence

Re-check, 27 September

  • 27 Sep 21:11: pgrep -f run_queue.sh inside an ssh command matched its own tailscaled be-child process twice; AGENT.md §11 still recommends pgrep -af queue.sh (D22)
  • the tailscaled journal logs every ssh command line (1,101 of our sessions on aifoundry1 since 25 Sep); aifoundry3's admin account is in adm

Back to the table · Requests

D16 Our scripts/deploy-lab.sh used macOS tar flags and failed on Linux#

Fixed 25 Sep blocks work any Linux client, including the lab hosts

27 September Fixed 25 Sep Since 25 Sep (us)

Holds: scripts/deploy-lab.sh uses the portable tar wrapper.

Next: nothing left

The 25 September write-up, kept as history:

Who acts (25 Sep): Done

What a user sees

tar: unrecognized option --no-mac-metadata, and nothing is deployed.

What went wrong, and why

known The script was written on macOS and passed bsdtar-only flags (--no-xattrs --no-mac-metadata) to tar unconditionally; GNU tar rejects them.

How we found it

Deploying from aifoundry2.

What it cost

Deploys from a lab host had to fall back to rsync and hand-run cmake.

What was done

Fixed in the repository in commit de26273 (25 Sep 07:21): a tar_create wrapper passes those flags only when tar is bsdtar. deploy-lab-gpsdk.sh has had the same wrapper since d07b4e0 (18 Sep).

How to verify

25 Sep about 17:25 on aifoundry2 with GNU tar 1.35: the old line fails with unrecognized option --no-mac-metadata, the new wrapper streams workloads/sgemm.

How to roll back

git revert de26273 -- scripts/deploy-lab.sh.

What the Nekko team can do
  • Nothing for the lab.
Evidence

Re-check, 27 September

  • scripts/deploy-lab.sh uses the tar_create wrapper (l.18, l.26)

Back to the table · Requests

D17 A coding agent told not to touch the cards ran a real measurement block on a shared card#

Open workaround corrupts results one lab host

27 September Open workaround Since 25 Sep

Safeguards extended for the heat code; fail-closed dry runs still to do.

Safeguards extended 27 Sep for the heat-placement code, before any card time: a hard refusal of aifoundry1 card 0 at every entry point (run_queue.sh, hplib.sh, block.sh), V3_FORCE honoured only under V3_DRY, validation blocks refuse without the pre-registration hash; an independent verifier ran 99 forbidden combinations (every card-0 spelling) in dry runs and all were refused. The code-writing and fixing agents ran only V3_DRY=1 and touched no card. V3_DRY is still only an environment variable (the report's 'fails closed' is not done), and a code review first found card 0 not refused (fixed before deployment).

Next: Us U13; Nekko team PO2

The 25 September write-up, kept as history:

Who acts (25 Sep): Users running agents; the lab (banner)

What a user sees

Card time was used, and data written, by a process that was supposed to be only writing code.

What went wrong, and why

known A cd inside a backgrounded && chain did not apply, so the agent's supposed dry run ran a real block. Nothing stopped it: our dry-run mode was only an environment variable (V3_DRY=1).

How we found it

Our own review of the version-3 work.

What it cost

An unplanned card session on a shared machine.

The workaround

By us: agents that only write code must test with V3_DRY=1 from an absolute path, and et-who checks for unexpected holders; since A12 the lock files exist on every host and our blocks take them. Any user's coding agent could make the same mistake.

What the Nekko team can do
  • Users running agents: give them a dry-run mode that fails closed, take the card lock in every tool that opens a card, and check et-who.
  • The lab could say in the banner that agents follow the same rules as people.

Requests: PO2

Evidence

Re-check, 27 September

  • 27 Sep: the heat code refuses aifoundry1 card 0 at every entry point; an independent verifier's 99 forbidden combinations were all refused in dry runs; code-writing agents ran only V3_DRY=1
  • V3_DRY is still only an environment variable (not fail-closed); a code review first found card 0 not refused (fixed before deployment)

Back to the table · Requests

D18 Our published notes and the lab-access page still carry the old diagnoses#

Partly fixed wastes time documentation

27 September Partly fixed Since 25-27 Sep (repository and mirrored pages corrected)

The old diagnoses are corrected in the repository and the mirrored pages (25–27 Sep); the public accounts page's correction is prepared but not published; new errors moved to D23.

Done after the report: commits eb7c762, 801d329, 336c3ac and 39836c7 (54 fixes across the findings, getting-started and the published pages); docs/lab-access.md now has Tailscale check mode and the 404 fix, et-who, the card locks, the banner, PATH, core dumps and dmesg. Left: the public accounts page of 18 Sep, which the banners link, is older than lab-access.md. Its correction (PATH, et-who and the card lock, /tmp at boot, starting-password advice) is prepared (U7), but publishing it was not allowed in the 27 Sep session and waits for the owner's OK.

4 Oct: the public accounts page of 18 Sep is still live and unchanged, and the banners still link it; it contradicts the New user page that newcomers now follow. Publish U7, or retire the page and point the banners at the New user page (a banner change, then a root install like U28).

Next: Us U7

The 25 September write-up, kept as history:

Who acts (25 Sep): Us (the repository); Nekko (the lab's page)

What a user sees
What went wrong, and why

known Written before the 25 Sep investigation and fixes; the corrections were collected but not yet applied.

How we found it

Cross-checking the sources for this report.

What it cost

Anyone reading the public pages gets aifoundry1, aifoundry3 and the access rules wrong.

Where it stands

Not done as of 18:10: the repository (at 8ca8d20) still has the old wording.

What the Nekko team can do
  • We apply the corrections and redeploy the affected pages: aifoundry1's cause and fix, aifoundry3's boot-time pin, the runtime per host, the toolchain lesson, the four-card firmware table, and the access page.
  • Nekko links the lab's own onboarding page from the banner.
Evidence
  • validate3/lessons.md (CORRECTIONS from the aifoundry1 investigation)
  • fix plan, After the fixes

Re-check, 27 September

  • commits eb7c762, 801d329, 336c3ac, 39836c7 (54 fixes); docs/lab-access.md has check mode, the 404 fix, et-who, locks, banner, PATH, core dumps, dmesg
  • left: the public 18 Sep page aifoundry-lab-accounts (ours; linked from all three banners) is older than lab-access.md; new wrong statements moved to D23

Back to the table · Requests

D19 Two host-to-card copies on one stream move less than one, and the link gives no full-duplex gain#

New workaround wastes time all three measured cards

30 September New workaround Since 27 Sep (E50)

Narrowed on 29 Sep by E55 (pre-registered, three cards): only two copies on one stream collide, moving 0.49 of one, while one copy on each of two streams moves 1.01, so the loss is per stream; a shared DMA read engine and the IOMMU are refuted, and the IOMMU-passthrough boot (U26) is no longer needed. Both directions at once give 1.07–1.08× (E50). Every card and root port run MaxPayload 256 B and MaxReadReq 128 B (28 Sep).

28 Sep, read as root with lspci -vv (configuration space through sysfs; no card node opened) on all four cards and their root ports: MaxPayload 256 B, the largest either side supports, and MaxReadReq 128 B; links at 16 GT/s ×8, the cards' full width; ASPM off; extended tags on the cards, off on the root ports. A 128 B read request is a quarter of PCIe's 512 B default. Whether it limits a single copy's rate is not tested (DI5); E55 shows it does not cause the halving.

Next: Nekko team DI5

Who acts: Nekko (the DMA engine's documentation)

What a user sees

Two host-to-card DMA commands in flight on one stream move about half as much as one at a time (0.49× on all three cards). Two streams together equal one. Copies in both directions at once give only 1.07–1.08× the faster single direction.

What went wrong, and why

unknown The link trains at 16.0 GT/s ×8 on every card in every run; host-to-card reaches 12.46–12.60 GB/s (79–80% of 15.75) and card-to-host 10.41–10.54 GB/s, with a dip at 64 MB per copy. Whether the halving comes from the card's DMA engine, the IOMMU (all three hosts translate DMA addresses) or the host is not established.

How we found it

E50 (27 Sep 14:24–15:13 PDT), pre-registered, five runs per card; the pilot at 14:21 showed the halving first (predictions P3, P11 and P12).

What it cost

A program that overlaps copies to hide transfer time gets less bandwidth, not more; double-buffered input pipelines must serialise their copies.

Where it stands

Worked around: one copy command in flight per stream (the barrier flag on every copy).

What the Nekko team can do
  • Send the PCIe DMA engine's documentation, or the register values the firmware programs (channels, read-request size, element split): the hub's rung 29 (request DI5).
  • With the lab's agreement we can run the command-size sweep and an IOMMU-passthrough boot (rung 35).

Requests: DI5

Evidence

Root steps, 28 September (verified 08:00–08:02; the steps after that, by their own output)

  • lspci -vv, 28 Sep: aifoundry2 07:14, aifoundry3 07:26, aifoundry1 07:56 (both cards); re-read at 08:02 on aifoundry1 and aifoundry3 with the same values

Back to the table · Requests

D20 A small copy or an empty kernel costs hundreds of microseconds, mostly the runtime polling for completion#

New workaround wastes time all three hosts (the runtime)

27 September New workaround Since 27 Sep (E50)

E50: 556–566 µs for a waited empty kernel against 104 µs queued; 377–411 µs for a lone 4 KB copy.

Next: Nekko team RT5, RT2

Who acts: Nekko (runtime)

What a user sees

An empty kernel on 32 shires takes 556–566 µs from launch to completion when the program waits for it, but 103.6–104.1 µs per launch when 100 are queued. A lone 4 KB copy takes 377–411 µs. Back-to-back round trips lock onto one of two latencies, about 0.11 ms or about 0.58 ms at the same size, depending on the order of calls.

What went wrong, and why

inferred Most of the wait is the runtime polling for completion, with two fixed polling modes; the figures follow each host's runtime build (C18).

How we found it

E50 (27 Sep): predictions P8, P9a and P9b, and the pilots.

What it cost

Launch and synchronisation overhead dominates any kernel shorter than about half a millisecond, and host-side timings differ between the hosts' runtime builds.

Where it stands

Worked around: queue launches and copies, and wait once; compare host-side timings only within one host.

What the Nekko team can do
  • Interrupt-driven completion, or a short adaptive poll, in the runtime; document the two polling constants (request RT5).
  • One runtime build on every host (C18, request RT2).

Requests: RT5, RT2

Evidence

Back to the table · Requests

D21 Documents the open drop cites but does not contain, so parts of the chip must be inferred#

New workaround wastes time documentation (the open drop)

27 September New workaround Since 27 Sep

The chip diagram and the memory-level pages mark these parts “inferred”.

30 Sep: experiments on 28–29 Sep settled some parts the pages had marked inferred (E55–E57: host writes land in their line's L3 home; read replies cross the mesh y first), and the chip diagram and memory-level pages were updated. None of the cited documents has arrived; aifoundry2's DRAM map (PA[17] reads as a row conflict there) is one more thing they would settle (DI6).

Next: Nekko team DI6, CD1–CD6

Who acts: Nekko

What a user sees

The chip diagram, the memory-level pages and the energy model mark parts “inferred”, because the documents that would settle them are cited but not included.

What went wrong, and why

known Missing: the NoC reference manual and specification (cited in core-et and the Shire Description); the chip floorplan and the die's orientation; the minion shire's floorplan; the ET-SoC-1 Data Book and a final datasheet (its electrical and package thermal sections are “in a future release”, and there is no ball map); the dev card's schematic and board file (linked, not included); the memory shire's description or its programmed DDR registers; a way to read the NoC's router and bridge counters; and the design documents the drop cites (power, the microcontroller, the shire bus master, PLL and DLL initialisation, debug, the minion CSRs, DFT, Maxion). The storage cells are unsettled too: the lab lead said the chip does not use SRAM, while the Shire Cache Specification (pdf pages 48 and 51) and the datasheet describe SRAM macros; the minion L1 uses custom latch RAM.

How we found it

Drawing the chip diagram and its review (27 Sep), and starting the memory-level diagrams that evening.

What it cost

Every user re-derives the same parts, and the diagrams carry inferred parts.

Where it stands

Worked around: inferred parts are marked and cite the source used.

What the Nekko team can do
  • Send, or say which can be shared, in this order of value: the floorplan and die orientation, the memory shire and its DDR registers, the NoC manual, the Data Book or datasheet with thermal data, the card schematic, the shire floorplan, NoC register access, and the cited design documents (the hub's rungs 21–28; request DI6).
  • Settle the storage-cell question.
  • Answer the chip diagram's six short questions (section 2.7): which minion options the silicon has, whether each minion shire has its own supply, a die image the page may show, the die's numbers, a check of the parts drawn dashed, and what the vendors' terms let you state publicly.

Requests: DI6, CD1, CD2, CD3, CD4, CD5, CD6

Evidence

Back to the table · Requests

D22 Our own checklist tells agents to run pgrep -af queue.sh over ssh, which always matches itself#

Fixed 27 Sep wastes time our docs (AGENT.md)

30 September Fixed 27 Sep Since 27 Sep 23:34 (a745199)

Fixed on 27 Sep at 23:34 and holding: AGENT.md §11 step 4 uses pgrep -af '[q]ueue.sh' and says why (commit a745199, pushed that night), and our card-behaviour notes and public newcomer brief carry the bracketed form. The 27 Sep update counted it as not committed yet.

Next: nothing left

Who acts: Us

What a user sees

ssh aifoundry3 '… pgrep -f run_queue.sh …' finds a process even when no queue runs. On 27 Sep at 21:11 PDT a starter refused to start twice for that reason.

What went wrong, and why

known Tailscale SSH runs every remote command as tailscaled be-child ssh … --cmd=<the whole command>, so an unbracketed pattern matches the command's own process (D15). AGENT.md §11 step 4 still recommended pgrep -af queue.sh.

How we found it

Restarting the heat round on aifoundry3, 27 Sep 21:11.

What it cost

Agents conclude a queue is running when none is, or kill their own session with pkill -f.

Where it stands

Corrected in the working copy on 27 Sep (U6): AGENT.md's step 4 and 14-card-behaviour.md's traps now use a bracketed pattern (pgrep -af '[q]ueue.sh') and say why. Committed and pushed on 27 Sep at 23:34 (a745199). Our starters move to a pid file after the heat work (U13).

What the Nekko team can do

Nothing needed from the Nekko team; the onboarding draft (appendix A) carries the bracketed form.

Evidence
  • the session log, 27 Sep 21:11–21:12 (two refusals, then the start at 21:12:01)
  • AGENT.md §11 step 4 (the committed text)

Back to the table · Requests

D23 Our docs carry new wrong or sensitive statements: card 1 “firmware DVFS”, a root SSH route, a stale numpy note#

New wastes time our repository

30 September New Since 27 Sep

Partly corrected: card 1's rows in AGENT.md and our card-behaviour notes are right, but the lib.sh comment is not. Its correction was reverted on 28 Sep to keep lib.sh at the bytes our locked experiments use, so it still says aifoundry3 has no system numpy, and it waits for the owner's decision on that lock. add-lab-user.sh still names the root route; that is the owner's call (U14).

4 Oct: the lib.sh note is also wrong on the facts: aifoundry3 has had numpy 1.26.4 from Ubuntu since 25 Sep. The root route in add-lab-user.sh is now the lab's documented onboarding step (AS5), so only its wording is left (U14).

Next: Us U13, U14

Who acts: Us; the owner (the root route)

What a user sees

14-card-behaviour.md's card table and AGENT.md described aifoundry1's card 1 as governed (“firmware DVFS; idles at 600 MHz”); scripts/add-lab-user.sh line 4 says it needs root ssh to each machine; tools/claims-v3/lib.sh line 17 says aifoundry3 has no system numpy.

What went wrong, and why

known Card 1's clock never moved in 359,657 samples (C24); the root route is one AGENT.md keeps out of the public repository; aifoundry3 has had numpy since 25 Sep (H14).

How we found it

Checking our docs against the 27 Sep re-check.

What it cost

New agents trust the table and plan governor tests on card 1; the public repository names a root path.

Where it stands

Card 1's rows are corrected and pushed (27 Sep, U6); the lib.sh comment's correction was reverted on 28 Sep (8eb0e38) to keep lib.sh at its locked bytes, and waits for the owner's decision on that lock. Rewording add-lab-user.sh is the owner's call (U14).

What the Nekko team can do

Nothing needed from the Nekko team.

Evidence

Back to the table · Requests

D24 Our PCIe probe held a card's lock for 12.8–14.9 s per run, over the lab's 10 s rule#

Fixed 28 Sep cosmetic three cards, 27 Sep 14:24–15:13

28 September Fixed 28 Sep Since 28 Sep 23:56 (us, U13)

Fixed on 28 Sep: run_pcie.sh releases the lock between sub-tests, and its first card run (aifoundry1's card 1, 23:56) took at most 1.84 s per sub-test, with no lock wait; the stale copies on aifoundry1 and aifoundry3 were replaced.

Next: nothing left

Who acts: Us (fixed on 28 Sep)

What a user sees

Each pciebench run took the card lock for its six processes together, 12.8–14.9 s.

What went wrong, and why

known The run script takes the lock once around all its sub-tests. Each device-opening process stayed under 10 s (at most 2.1 s), which is how the rule is meant.

How we found it

Recorded in the experiment's README.

What it cost

Minor.

Where it stands

AGENT.md now says the rule applies per device-opening process, with the lock released between sub-tests (working copy, U6). run_pcie.sh will release the lock between sub-tests before any rerun (U13). Fixed on 28 Sep: the script releases the lock between sub-tests (commit 031c723, dry-tested that morning). The hosts still had the copy from before the fix, which a read-only audit found at about 21:00; the copies on aifoundry1 and aifoundry3 were replaced, and at 01:10 on 29 Sep both match the repository's (sha256 9dbf1d80…). Its first card run, on aifoundry1's card 1 at 23:56, released the lock between sub-tests and never waited for it; every sub-test exited 0, and the longest (bandwidth) took 1.84 s.

What the Nekko team can do

Nothing needed from the Nekko team.

Evidence

The evening of 28 September

  • docs/reports/data/2026-09-29-pcie2/run_pcie-r101/: run.json (lock_released_between true, lock_waits 0, every exit 0; per sub-test 0.18–1.84 s)
  • 29 Sep 01:10: sha256sum ~/nekko/workloads/pciebench/run_pcie.sh on aifoundry1 and aifoundry3, equal to the repository's

Back to the table · Requests

D25 Reading the host CPU's energy needs root, so CPU-against-card energy comparisons rest on an assumed CPU power#

New wastes time all three hosts

4 October New Since 29 Sep (found)

The host CPU's energy counter (/sys/class/powercap/intel-rapl:0/energy_uj) is readable by root only on all three hosts (checked 4 Oct), the kernel's default since a 2020 side-channel fix (CVE-2020-8694). So our 29 Sep comparison of CPU and card energy had to assume 125–251 W for the CPU. A policy choice, not a fault: the lab decides (MO8).

Next: Nekko team MO8; then us (root)

Who acts: Roman (the decision); us, with root

What a user sees

cat /sys/class/powercap/intel-rapl:0/energy_uj fails with “Permission denied” for an ordinary user.

What went wrong, and why
  • known The counter is mode 0400 root, the kernel's default since CVE-2020-8694 (PLATYPUS) showed that it can leak information across users.
How we found it

Our sparse-parity energy comparison of 29 Sep.

What it cost

That comparison rests on an assumed 125–251 W for the CPU, so its energy ratio has a wide range.

Where it stands

Reports state the assumption. Three ways out, for the lab: keep it root-only; give a group read access with a udev or tmpfiles rule, accepting the side channel; or have a small root service publish package energy coarsely, for example once a second, which removes the side channel (MO8).

What the Nekko team can do

Requests: MO8

Evidence
  • ls -l /sys/class/powercap/intel-rapl:0/energy_uj on all three hosts, 4 Oct

Back to the table · Requests

4. What changed on the machines#

4.1 On 25 September#

Everything below was done as root over Tailscale SSH, at the owner's request, after a read-only audit and a written plan; the scripted session was dry-run on every host first. No card node was opened as root, steps that touch a card or reboot (group B) ran with the cards idle, and no other user's processes were stopped (the tailscale upgrade briefly dropped SSH connections; the other user on aifoundry1 was still logged in afterwards). Every file the scripts changed was copied first to /root/labfix-20260925/replaced/, and each machine keeps the session log (A-session-*.log), the window logs (W*.log) and the upgrade log (B3-upgrade.log) in /root/labfix-20260925/. The aifoundry1 driver fix has its own fix log. On 27 September every one of these changes was still in place on every host where it was applied.

aifoundry1#

Not rebooted. The driver fix and the kernel-update repair ran at 15:00–15:11, the scripted session at 16:24–16:27, the upgrade window at 16:27–16:39, and the banner warning about card 0 at 17:56.

Time (PDT)ChangeBackupRoll backVerified
15:01–15:02Driver rebuilt by DKMS from the fixed 353f20e source for 7.0.0-30, -31 and -34 and reloaded; /etc/modules-load.d/et_soc1.conf added C1, C2/root/et-soc1-fix-20260925/ (old source, .ko files, modinfo)Section 1 of the troubleshooting report0.20.0 / 47D26A30…; firmware query on both cards; queues on both cards since 17:13
15:03, 15:11Caches cleared: apt cache, archived journal, ten disabled snap revisions, root's Hugging Face chunk cache and ccache H3none (re-created on demand)not applicable6.4 GB free afterwards
15:03:40dpkg --configure -a finished the 7.0.0-34 kernel update H5nonenot needed; 7.0.0-31 stays installeddpkg --audit empty; the initramfs has ZFS
16:24:32A1: backups; ZFS snapshots @labfix-20260925-A/root/labfix-20260925/destroy the snapshots after 27 Seplisting in the log
16:24:34A8: log at most one corrected AER error a minute for card 0's root port (labfix-aer-ratelimit-et0.service) C16aer-ratelimit.beforedisable and remove the unit; write 5000 and 10 back60000 1; kern.log flat; RxErr still counting
16:24:34A9: persistent journal, 1 GB cap, 1 GB kept free H9/etc tarballremove the drop-in; restart journaldsettings shown; 51 MB at 17:19
16:24:34–16:25:05A2, A3, A14a: purged kernel 7.0.0-30, headers of kernels not installed, openipmi, rc packages; removed the esperanto DKMS driver, orphan /lib/modules directories and old runner versions C3, H3dpkg selections; usr-src-esperanto-0.20.0.tgz, var-lib-dkms-esperanto.tgz, lib-modules-orphans.tgzreinstall the packages; untar, dkms add, dkms installdkms status: et-soc1 only; /lib/modules: 7.0.0-31, -34
16:25:05–16:26:21A4, A5: numpy, venv, scipy, matplotlib, pandas, build libraries and tools; one pending security update H14dpkg selectionsapt-get purge the listnumpy 1.26.4, venv ok
16:26:24–16:26:31A6: chrony replaces timesyncd. A7: systemd-coredump and an unlimited soft core limit H10, H16timesync status; sysctl -ainstall timesyncd; install apport-core-dump-handler, remove the two filesNormal, 8 sources; core_pattern
16:26:39A12: et-who, banner, PATH, lock files for both cards, et-lab-manifest. A13: ET crash reports moved out of /var/crash C11, C19replaced/; var-crash/remove the files (C11)et-who as nobody; locks 0666
16:26:39–16:26:44A10 plymouth quit; A11 cleared failed units; A16 Wi-Fi power saving off; A15 performance profile H8, H11, H17power-profile filepowerprofilesctl set balanced; remove the NetworkManager drop-inrunning; performance; power save off
16:27:43B2: GRUB menu shown for 5 s. C1.1: dmesg and perf for users H21, H15/etc tarballremove the two files; update-grub; sysctl -w the old valuesset timeout=5; dmesg as a user
16:27–16:39B3: full upgrade (200 packages) with Docker held, snapshots @labfix-20260925-B first, the CI runner stopped and restarted H18snapshots B; upgradable.pre-B3.txtroll back to snapshots B from a rescue shell; apt-mark unholdResult=success; dpkg --audit empty
17:56:40A warning in the login banner (/etc/motd) that card 0 overheats under load C21motd.before-card0-warningcopy the backup back to /etc/motdbanner shows it (18:07)

aifoundry2#

Not rebooted (it ends our own session and wipes /tmp). Root reached through aifoundry3.

Time (PDT)ChangeBackupRoll backVerified
16:23:02A1: backups/root/labfix-20260925/not applicablelisting in the log
16:23:03–16:23:13A2, A3: purged headers of kernels not installed and rc packages; removed 22 stale build products from the DKMS source tree and orphan /lib/modules directories C3dpkg selections; et-soc1-src-build-products.tgz; lib-modules-orphans.tgzreinstall; untardkms status clean
16:23:13–16:24:03A4: venv, scipy, matplotlib, pandas and tools (not Ubuntu's tmux, H25) H14dpkg selectionsapt-get purge the listnumpy 1.26.4, venv ok
16:24:03–16:24:09A6 chrony; A7 systemd-coredump H10, H16timesync status; sysctl -aas on aifoundry1Normal, 8 sources; core_pattern
16:24:16A9 journal cap 2 GB; A12 onboarding; A13 one ET crash report moved; A14b /etc/hosts names the machine, tmux snap held H9, C11, H23, H25replaced/remove the files; restore /etc/hosts from replaced/; snap refresh --unhold tmux127.0.1.1 aifoundry2; tmux held
16:24:16–16:24:21A10 plymouth quit; A16 Wi-Fi power saving off (file; applied later); A15 performance profile H8, H11, H17power-profile fileas on aifoundry1running; performance; power save off by 17:19
16:28:02B2 GRUB menu; C1.1 dmesg and perf for users H21, H15/etc tarballas on aifoundry1set timeout=5; dmesg as a user
16:28–16:36B3: full upgrade (212 packages), the CI runner stopped and restarted H18upgradable.pre-B3.txtimpractical (SRUs)Result=success; dpkg --audit empty

aifoundry3#

The only machine rebooted, and the only one where the per-card reset was tried.

Time (PDT)ChangeBackupRoll backVerified
16:14:01A1: backups/root/labfix-20260925/not applicablelisting in the log
16:14:02–16:14:10A2, A3: purged headers of kernels not installed and rc packages; removed the hand-swapped 6.17 module and orphan /lib/modules directories C3dpkg selections; lib-modules-6.17.0-40-updates-dkms/; lib-modules-orphans.tgzreinstall; untardkms status clean
16:14:10–16:14:50A4: numpy, venv, scipy, matplotlib, pandas, tmux, htop, smartmontools and tools H14dpkg selectionsapt-get purge the listnumpy 1.26.4, venv ok
16:14:50–16:15:07A6 chrony; A7 systemd-coredump; A9 journal cap 2 GB; A12 onboarding; A10 plymouth quit; A16 Wi-Fi power saving off; A15 performance profile H10, H16, H9, C11, H8, H11, H17as on aifoundry1as on aifoundry1VERIFY block at 16:15:07
16:15:36B1: driver loaded from modules-load.d and before the clock guard. B2: GRUB menu. B6: runtime backups moved from /opt/et/lib to /var/backups/aifoundry3-opt-et-20260723 C2, H21, C18/etc tarballremove the two B1 files and daemon-reload; remove the GRUB file and update-grub; move the backups backafter the reboot: driver at +6.2 s, guard on attempt 1; left 0, moved 5
by 16:19:59B3: full upgrade (216 packages) H18upgradable.pre-B3.txtimpractical (SRUs)Result=success
16:20:36B4: reboot into 7.0.0-34 (up at 16:21:05) H67.0.0-31 kept installedchoose 7.0.0-31 in the GRUB menu0.20.0; guard marker = boot_id; settings survived
16:22–16:22:29C1.1 dmesg and perf for users; plymouth quit again; C1.2 one per-card reset of the idle card, then the guard restarted H15, H8, C14, C6nonenot applicablenodes back 0666; later runs worked; 0 W / 600 MHz at 17:17

4.2 On 27 September#

Everything ran as our ordinary user between 22:10 and 22:28 PDT. No root command ran: this session's permission check refused root over ssh (the root wrapper at 22:17, and a read-only root login at 22:18), so nothing was installed that night, and the steps that need root or the owner's OK waited (the second table; most of them ran on 28 September, section 4.3). No card node was opened, no process was stopped, and the running experiments were not disturbed: aifoundry1's card-1 queue had started at 22:13, so only file-only steps ran there, and aifoundry3's heat queue had ended at 22:11. The read-only re-check before it (21:26–21:45) changed nothing on the machines, nor did a read-only check as the user at 23:00–23:30, which corrected H23, C2, C7, H3, H31, H32 on this page; the copy of our /tmp files was re-synced at 22:21 and 23:29 (U1). Our steps are named U1–U14 (section 4.4).

aifoundry1#

Time (PDT)ChangeBackupRoll backVerified
22:15U4: moved our own crash report _opt_et_bin_dev_mngt_service.1009.crash (from our aborts of 26 Sep) out of /var/crash to ~/nekko/labfix2/var-crash/, as the user. File-only: card 1's queue had started at 22:13 C19moved, not deleted (sha256 unchanged)move it back/var/crash is empty
22:18U10, part: the memory table that udev exports, the IOMMU groups and the CPU, as the user, read-only. No card sysfs reads while card 1's queue ran H29not needed (read-only)not neededfour 32 GB DDR4-3200 DIMMs on two channels; IOMMU translating (DMA-FQ)
by 22:19Staged in ~/nekko/labfix2/, as the user: the 27 Sep et-who, et-holders, et-lab-health (rev 2), this host's banner, the install and rollback scripts (U2, U5, U8) and the snapshot script (U3). Nothing installed H30, C12, C24, C25not needed (nothing replaced)rm -rf ~/nekko/labfix2checksums match the install script

aifoundry2#

Time (PDT)ChangeBackupRoll backVerified
22:12U1: copied /tmp/claude-1019 (11 GB: this page's generator and data, the 25 Sep audit and fix records, the firmware research) to ~/claude/private/tmp-claude-1019-20260927/ (mode 700), as the user, with nice and ionice; re-synced at 22:21 and 23:29. Work continues in /tmp, so it is copied again before any reboot H22not needed (a copy)rm -rf the copyrsync exit 0 in 10.4 s; a dry run afterwards lists only files written since. The copy is 231 MB larger, because rsync without -H stores the 262 MB of hard-linked files twice
22:18U10, the part a user can read: the memory table, the card's link speed and width, the IOMMU group types and the kernel command line, read-only H29not needed (read-only)not neededtwo 32 GB DDR4-2666 DIMMs on two channels; the card's link 16 GT/s ×8; IOMMU translating (DMA-FQ)
22:19Staged in ~/nekko/labfix2/, as the user: the same tools, this host's banner and the scripts (U2, U5, U8). Nothing installed H30, C12, C25not needed (nothing replaced)rm -rf ~/nekko/labfix2checksums match; et-lab-health run from the stage as a user: exit 0, no warning

aifoundry3#

Time (PDT)ChangeBackupRoll backVerified
22:17U9, part: saved the world-readable wtmp reboot history (last -x -F reboot shutdown) to ~/nekko/labhealth/aifoundry3-boots-20260927/, and a copy to ~/claude/private/ on aifoundry2, as the user, after the heat queue R3 had ended at 22:11 H32, H7not neededrm -rf the directory137 lines back to 2 Jan: 125 boots, 10 clean shutdown records
22:18U10, the part a user can read: the memory table, the card's link, the IOMMU groups H29not needed (read-only)not neededone 32 GB DDR4-2666 DIMM, in ChannelA-DIMM1, three slots empty: single-channel memory; the card's link 16 GT/s ×8
22:19Staged in ~/nekko/labfix2/, as the user: the same tools, this host's banner and the scripts (U2, U5, U8), plus the journal export (U9). Nothing installed H30, C12, C23, H31not needed (nothing replaced)rm -rf ~/nekko/labfix2checksums match; et-lab-health run from the stage as a user: exit 0, no warning

The repository (no machine)#

Time (PDT)ChangeBackupRoll backVerified
by 22:26U6: our docs corrected in the working copy: card 1's clock and card 0's governor (AGENT.md, 14-card-behaviour.md), the bracketed pgrep, the 10 s rule per process, et-who --check, et-lab-health and “/tmp is cleared at boot” in lab-access.md, the lib.sh comment, aifoundry3's memory in E50's hosts.txt, and a new tools/lab/ with the tools as installed and staged D22, D23, D24, H30, C24gitgit checkout the filesnot committed yet
by 22:28U12: drafts of the onboarding page and the per-card sheet, for the lab to adopt (requests DI1, DI2) C4, C11, H1, H22not needednot applicableappendices A and B of this page

Not done on 27 September, and why#

StepWhereWhy notWhat it needs
Install the new et-who, the corrected banners and et-lab-health (U2, U5, U8)all threethis session's permission check refused root over ssh (the root wrapper at 22:17:19, and a read-only root login at 22:18:33)root, once the owner allows it, or the owner runs the staged script; aifoundry1 only in its 02:05–05:45 gap or after 08:00 on 28 Sep. 28 Sep: done on all three (section 4.3)
Urgent: export aifoundry3's July journal and raise our journal cap to 4 GB (U9, root part)aifoundry3root refusedroot, as soon as the owner allows it: the journal is full at our 2 GB cap and deletes the 17 Jul files next (H32). 28 Sep: done, 06:44 and 07:11 (section 4.3)
Read MaxPayload and MaxReadReq (U10, root part)all threeroot refusedroot; aifoundry1's card links in its next gap. 28 Sep: done on all three (section 4.3)
Stop our two orphaned headless Chrome trees (U11)aifoundry2refused twice as interfering with workloads, although both trees are ours, orphaned and without connectionsthe owner's OK: kill -TERM the two tree roots after re-checking them. 28 Sep: refused again; three trees then, stopped at 08:34 with the owner's approval (section 4.3)
Mask the desktop portal and the notifier in our own user managers (U11)aifoundry2 (then the other two)refused as persistencethe owner's OK. 28 Sep: the portal masked on aifoundry2; the notifier and the other two hosts not yet
Publish the accounts runbook correction (U7)spacesheeprefused as a write to an external systemthe owner's OK
Destroy our 25 Sep ZFS snapshots (U3)aifoundry1not due: the report said after 27 Sep, and it needs rootthe 02:05–05:45 gap on 28 Sep or after 08:00, with root. 28 Sep: dry run only at 07:53, the destroy refused; destroyed at 08:33 with the owner's approval (section 4.3)

4.3 On 28 and 30 September (root, at the owner's request)#

At 06:30 PDT the owner gave us root again (“Fix any issues that require root”). Everything below ran as root over Tailscale SSH (aifoundry2 through aifoundry3), from the written plan of 27 September, with each script dry-run first; the tools were installed exactly as the README of the repository's tools/lab/ says, after checking that plain et-who still prints the same holder lines. None of the steps of 06:44–07:57 opened, queried or reset a card, and none touched the driver, DKMS, /opt/et, the firmware, aifoundry3's 600 MHz boot service, the network, SSH, accounts or sudo policy (our 25 September et-who sudoers drop-in is unchanged); no machine was rebooted and no one else's process was stopped. aifoundry1's steps waited until card 1's validation queue had ended (07:51). Every file replaced was copied first to /root/labfix-20260928/replaced/, and each machine keeps its session log (/root/labfix-20260928/session-*.log); our copies of the commands and outputs are in labreport2/fix28-aifoundry*.txt. A separate read-only check at 08:00–08:02 confirmed every change of 06:44–07:57 below; it opened no card node and changed nothing, apart from two 2-second holds of card 0's lock file on aifoundry2 and aifoundry3 to test et-who --check. aifoundry2's table also has the two resets of its hung card (C27), with the owner's approval: the per-card sysfs reset at 06:39:27, before these steps, which did not recover the Master Minion, and the management reset at 08:32:45, which did (their logs are in /root/labfix-20260928/ on aifoundry2). The two steps the permission check had refused, U3 and the Chrome part of U11, ran at 08:33 and 08:34 with the owner's approval; they came after the check, so their rows give their own output as the evidence.

aifoundry3#

Time (PDT)ChangeBackupRoll backVerified
06:44U9, export first: copied the whole persistent journal (102 files, 2,239,758,336 B: 48 boots, from 17 Jul 01:14) to /root/labfix-20260928/journal-export-20260928/ (mode 0700), before touching the cap H32, H7the export is the copyrm -rf the exportthe same file count and bytes as the live journal, and journalctl -D lists the same 48 boots; at 07:08 every archived file was byte-identical to its copy. One archived file of 23 Jul fails journalctl --verify (“Bad message”), live and copy alike
06:50U9: each boot's last kernel and PID-1 lines (no user command lines) to ~/nekko/labhealth/aifoundry3-boots-20260927/, copied to ~/claude/private/ on aifoundry2 at 07:27 H7, H32not needed (new files)rm -rf the directory100 files on each side, from the 17 Jul 01:14 boot
07:11:29U9, the cap: SystemMaxUse from 2G to 4G in our /etc/systemd/journald.conf.d/60-labfix.conf, then a journald restart H32, H9replaced/etc/systemd/journald.conf.d/60-labfix.confcopy it back; restart journald“System Journal … is 1.8G, max 4.0G, 2.1G free”; at 08:00 journald active since 07:11:29 with 0 restarts, nothing vacuumed after 07:11, 46 boots kept from 17 Jul 01:21. Between the export and the cap, journald had deleted its two oldest files (the 17 Jul 01:14 and 01:18 boots): they are only in the export
07:26:01U2, U5, U8: installed the new et-who (with --check), et-holders (the new idle sentence), et-lab-health and this host's banner from the repository's tools/lab/, as its README says: dry run at 07:25:41, each file to a .new name and then mv. et-lab-manifest, the banner script and the sudoers drop-in are unchanged H30, C11, C12, H22, C23, H31replaced/usr/local/bin/et-who, replaced/usr/local/sbin/et-holders, replaced/etc/motdcopy the three back; rm /usr/local/bin/et-lab-healthsha256 equal to tools/lab/; sudoers parses; as nobody and as the user, et-who exits 0, --check 0 when free and 1 when card 0's lock is held (08:01:48, 2 s of flock -n), a bad argument 2; the frozen experiment code's parse finds the same holder lines; et-lab-health exits 0, no WARN
07:26:36U10, root part, read-only: lspci -vv of the card and its root port (configuration space through sysfs; no card node opened) and dmidecode memory D19, H29not needed (read-only)not neededMaxPayload 256 B, MaxReadReq 128 B on both; link 16 GT/s ×8; one 32 GB DDR4-2666 DIMM (ChannelA-DIMM1), three slots empty. Re-read at 08:02: the same

aifoundry2#

Time (PDT)ChangeBackupRoll backVerified
06:39:27The per-card sysfs reset of the hung card (C27), with the owner's approval, before the steps below: echo 1 > /sys/bus/pci/devices/0000:02:00.0/soc_reset/reinitiate C27not applicablenot applicablethe kernel log: enabling device, added peer-to-peer DMA memory. But our test launches at 06:41, 06:43 and 06:47 still failed (exit status 1; the two saved, at 06:43 and 06:47, with Couldn't use the HPSQ. Perhaps the Master Minion is hanged?): it did not recover the Master Minion
07:09:56U2, U5, U8: the same tools and this host's banner, installed the same way (dry run at 07:08:31), through aifoundry3 H30, C11, C12, H22, H28replaced/usr/local/bin/et-who, replaced/usr/local/sbin/et-holders, replaced/etc/motdas on aifoundry3sha256 equal to tools/lab/; sudoers parses; et-who exits 0, --check 0 free and 1 held (07:10:40 and 08:01:34), a bad argument 2, as the user and as nobody; et-lab-health exits 0, no WARN
07:14:30U10, root part, read-only, as on aifoundry3 D19, H29not needed (read-only)not neededMaxPayload 256 B, MaxReadReq 128 B on the card and its root port; link 16 GT/s ×8; two 32 GB DDR4-2666 DIMMs, one per channel
07:28:33U11, part, as our user: masked xdg-desktop-portal and its GTK backend in our own user manager (about 260 failures a day each, started by our screenshot runs' Chrome), and cleared their failed state at 07:29:46 H33, H26none needed: two links to /dev/null in ~/.config/systemd/usersystemctl --user unmask the twoboth masked; the user manager is “degraded” only by the firmware notifier, which was failing before
08:32:45U15: the management reset of the same card, with the owner's approval, as our user after et-who showed no holder: dev_mngt_service -m DM_CMD_RESET_ETSOC -n 0 C27, C14not applicablenot applicablethe kernel log: Mgmt: Device is resetting, then the card re-enabled and its DMA memory re-added; “Service request succeeded” at 08:32:52. At 08:33 a 1 s test kernel ran 3 launches, all ok at 600 MHz (labreport2/fix28/c27/, labreport2/fix28-aifoundry2.txt)
08:34:30U11, its Chrome part, with the owner's approval: stopped our three orphaned headless Chrome processes (20 h and 10 h old, parent PID 1) and their children with a plain kill H26none needed (our own processes, with no client connected)not applicableat 08:34:44 none of the three was left; at 08:36:59 and 08:52:37 ps -u lists no Chrome process of ours (labreport2/fix28-aifoundry2.txt)

aifoundry1#

Time (PDT)ChangeBackupRoll backVerified
07:54:51U2, U5, U8: the same tools and this host's banner, installed the same way (dry run at 07:54:20), after card 1's validation queue had ended (we waited from 07:20) H30, C11, C12, C13, C21, C24, H22replaced/usr/local/bin/et-who, replaced/usr/local/sbin/et-holders, replaced/etc/motdas on aifoundry3sha256 equal to tools/lab/; sudoers parses; as nobody et-who exits 0, --check 0, a bad argument 2; the held test was skipped because another account was logged in. et-lab-health exits 1 on real warnings: card 0's power and thermal events and its root port's corrected errors (C21, C15), /home at 99% and the pool at 95% (H3)
07:55:30systemctl daemon-reload, which re-reads unit files and restarts nothing: 294 units, our AER unit among them, were flagged as changed on disk since the 25 Sep upgrade C16the unit lists before and after, in /root/labfix-20260928/not needed (a reload changes no unit)0 units need a reload, 0 failed, no unit changed state; the AER rate limit still 60000 ms, burst 1
07:56:20Compressed the flood-era kern.log.1 and syslog.1 (287 and 299 MB, rotated 27 Sep) with gzip -k, checked each .gz against the original, then removed the original C16the .gz holds the same bytes; the originals' sha256 in logs-sha256-before-gzip.txtgunzip the twogzip -t passes; each decompressed sha256 equals the original's (6f20f219…, 86e0696e…); owner syslog:adm, mode 0640, mtime kept; 11.7 and 13.6 MB; logrotate -d rc 0, and the next rotation renames kern.log.1.gz to .2.gz
07:56:50U10, root part, read-only, both cards D19not needed (read-only)not neededMaxPayload 256 B, MaxReadReq 128 B on both cards and both root ports; both links 16 GT/s ×8. Re-read at 08:02: the same
08:33:40U3, with the owner's approval: destroyed our 34 snapshots of 25 Sep (@labfix-20260925-A and -B, recursive, on rpool and bpool), after the dry run of 07:53 had listed only ours H3, H4none: they were the backup of 25 Sepnone: a destroyed snapshot cannot come backno labfix snapshot left; available space rpool 5.70 → 7.98 GB, bpool 1.36 → 1.65 GB; zpool status -v lists only the four files of H4 (labreport2/fix28/a1/s01-real.out)

Not done in the 06:44–07:57 run, and what became of it#

StepWhereWhy notWhat it needs, or what happened
Destroy our 25 Sep ZFS snapshots (U3)aifoundry1the dry run at 07:53 listed 34 snapshots, all ours (17 each of @labfix-20260925-A and -B), no holds and no other snapshot; the destroy itself was refused by this session's permission check at 07:54, so nothing was destroyed then. At 08:00 they held 1.93 GB on rpool/ROOT and 276 MB on bpoolDone at 08:33:40, with the owner's approval (aifoundry1's table above)
Stop our three orphaned headless Chrome trees (U11)aifoundry2refused by this session's permission check as interfering with workloads, although all three are ours (24 processes, from screenshot runs on 27 Sep) and no client was connectedDone at 08:34:30, with the owner's approval (aifoundry2's table above)
Stop apport's duplicate crash reportsaifoundry1not changed: the three apport units the plan named are already inactive here; the .crash copies come from apport's systemd-coredump hook, and switching that off was then the lab's crash-reporting decisionDone on 28 Sep at 20:51, with the owner's approval (U20; the evening's table)
A 200 MB size cap for rsyslog's logsaifoundry1left off on purpose, as in the 25 Sep plan: /etc/logrotate.d/rsyslog is rsyslog's own file, a local edit would stall its unattended updates, and the flood is rate-limited since 25 Sepnothing; the edit is staged, unused, as /root/labfix-20260928/rsyslog.logrotate.new
Mask the firmware notifier in our own user managers (U11)all threenot part of this run: only the portal on aifoundry2 was maskedDone at 21:20 as our user, on all three hosts, with wireplumber (U11)

aifoundry1, 28 September evening: four owner-approved items#

At 20:51 PDT, with the owner's approval, one script ran U27, U20, U22 and U24 on aifoundry1 (20:51:12–20:51:15), after a read-only look at each item's state at 20:49. It opened no card node, rebooted nothing and stopped no user's process; the host keeps its session log and the backups, and our copies of the script and its output are in fix28b/. A read-only check at 01:10–01:18 on 29 September found every change in place. The same items on aifoundry2 and aifoundry3 did not run (section 2.9).

Time (PDT)ChangeBackupRoll backVerified
20:51:12U27: the default target set to multi-user.target (it was graphical.target); it takes effect at the next boot H8, H33the old default, kept on the hostsystemctl set-default graphical.targetsystemctl get-default: multi-user.target (20:51 and 01:10)
20:51U20: bluetooth and cups-browsed disabled and stopped; the firmware-updater snap disabled; apport's drop-in on systemd-coredump linked to /dev/null, then systemctl daemon-reload H33, C19, H16not needed (unit and snap states, and one link)systemctl enable --now the two; snap enable firmware-updater; remove the link; daemon-reloadboth services inactive (20:51) and disabled (01:10); the snap disabled; the link in place; /var/crash empty (01:10)
20:51U22: one login-service setting (the command in section 2.9), then a reload H19kept on the hostremove the setting; reload the servicethe setting in effect (20:51) and in place (01:10)
20:51:15U24: a marked block in /etc/hosts with aifoundry2 and aifoundry3 at their tailnet addresses H23the old file, kept on the hostcopy it backgetent hosts: both at their tailnet addresses (20:51 and 01:18)

30 September, 14:37–14:41: the owner's root session

The owner ran our commands for the section 2.9 items that wait on no one else, as root, host by host; our own session's permission check had refused to run root commands. We checked the results read-only as our user at 14:41–14:50 (file times, systemctl, getent, an ssh attempt without credentials). The commands kept their backups under /root/labfix-20260930/ on each host.

TimeChangeRoll backVerified
14:37:52–54aifoundry2: U27 the headless target; U20 bluetooth, cups-browsed, the firmware-updater snap and apport's hook off; U22 password logins off; U24 the peer namesas in section 2.9get-default multi-user.target; both services disabled and inactive; the snap disabled; the hook's link in place; public-key logins only (14:40:27); both peers at their tailnet addresses
—aifoundry2: U16 and U18 not run—no et-reset, et-lab-health rev 2, no timer (14:48)
by 14:39:54aifoundry1: U16 et-reset; U18 et-lab-health rev 3 and its daily timer (the 28 Sep items U20, U22, U24 and U27 were already in place)remove et-reset; disable the timer, remove its units, reinstall rev 2et-reset, rev 3 and the timer, read at 14:39:54
14:39:50–52aifoundry3: U27; U20; U24, whose second line names aifoundry3 itself instead of aifoundry2as in section 2.9get-default multi-user.target; services and snap off; the hook masked; aifoundry2 still at its LAN address
14:39:59aifoundry3: U16 et-reset; U18 et-lab-health rev 3 and its daily timeras aboveet-reset root:sudo 0750; rev 3; first run ok at 14:40:00; the system running
14:40:21–14:41:21aifoundry1: U25, two readings at 16 GT/s, then the retrain of card 0's root port to 8 GT/sa reboot restores 16 GT/sthe host stopped answering within twenty seconds (C28)
about 15:07all three: power-cycled on site by Roman—all three on 7.0.0-34, running, every card's nodes present and links at 16 GT/s x8 (15:18–15:20)

4.4 Our own steps#

What does not need the Nekko team, we do ourselves: our tools, scripts and docs, the banner text we installed on 25 September, our processes, files and snapshots, read-only checks, drafts of the documents this report asked the lab to write, and, since the owner confirmed root on the three machines on 28 September, everything that needs only root there and no one else's decision (U15–U27, with their commands and risks in section 2.9). Their state on 30 September, from the read-only re-check at 14:17–14:33 PDT and our checks after the owner's root session (14:41–14:50):

StepWhatWhereState, 30 SepProblems
U1Copy our working files out of aifoundry2's /tmpaifoundry2Too late Copied 27 Sep 22:12, last re-synced at about 06:37 on 28 Sep; the reboot of 30 Sep 15:07 cleared /tmp, and the 19 GB written since were lost (H22)C8, C22, H6, H22
U2et-who --check, and et-holders' idle sentenceall threeDone 28 Sep, as root: aifoundry2 07:09, aifoundry3 07:26, aifoundry1 07:54; verified 08:00–08:02 (free and held cases)C11, H30
U3Destroy our 25 Sep ZFS snapshots on aifoundry1 (section 2.9)aifoundry1Done 28 Sep 08:33:40, with the owner's approval: all 34 snapshots destroyed, about 2.4 GB freed (rpool 7.98 GB available)H3, H4
U4Move our crash report out of aifoundry1's /var/crashaifoundry1Done 27 Sep 22:15. A second crash report of ours (wireplumber's, 28 Sep 10:27) removed 28 Sep 21:20; apport's hook, which writes them, is off on aifoundry1 since 20:51 (U20)C19
U5Corrected login banners (card facts, the commands that change a card, /tmp at boot)all threeDone 28 Sep, with U2 (same times); each banner matched tools/lab/ then; aifoundry1's and aifoundry2's have not since 2 Oct, when the repository's were corrected (U28)C5, C7, C12, C13, C17, C21, C23, C24, H22, H28, H31, D17
U6Correct our repository docs (card 1, the bracketed pgrep, the 10 s rule, lib.sh, lab-access.md, tools/lab/)repositoryPartly done Committed and pushed; the lib.sh comment was reverted on 28 Sep to keep lib.sh at its locked bytes (D23)C7, C13, C24, H30, D15, D18, D22, D23, D24
U7Update our public accounts runbookspacesheepWaits for the owner's OK Prepared; publishing needs no root, only the owner's OKH19, D18, D23
U8et-lab-health, a read-only health checkall threeDone Done 28 Sep (rev 2); rev 3 installed on aifoundry1 and aifoundry3 on 30 Sep with U18; aifoundry2 got rev 3 at 15:36 that day; all three run the repository's rev 3 (4 Oct)C1, C3, C6, C15, C21, C25, H5, H11
U9Save aifoundry3's boot history: export the journal's July boots and raise our 25 Sep journal cap from 2 GB to 4 GBaifoundry3Done wtmp saved 27 Sep 22:17; the journal exported 28 Sep 06:44 (48 boots from 17 Jul 01:14), per-boot tails 06:50, cap 4 GB 07:11H7, H9, H32
U10The hub's rung-30 readings (memory, PCIe payload sizes, IOMMU)all threeDone Memory, links and IOMMU as a user 27 Sep 22:18; payload sizes and memory as root 28 Sep (aifoundry2 07:14, aifoundry3 07:26, aifoundry1 07:56)H29, D19
U11Our background load: the orphaned Chrome, the desktop portal and notifier in our user managersaifoundry2 (all three)Done Our Chrome on aifoundry2 stopped 28 Sep 08:34 (with the owner's approval); the portal masked there at 07:28; the firmware notifier and wireplumber masked in our user managers on all three hosts at 21:20 (a mask stops restarts only: our wireplumber still runs on aifoundry2 and aifoundry3)H26, H33
U12Drafts of the onboarding page and the per-card sheetthis pageDone 27 Sep: appendices A and BC4, C5, C10, C23, H1, H28, H31
U13After the heat work: rebuild aifoundry1's host programs, the card-0 guard's lock, the PCIe probe's lock, less polling, fail-closed dry runs, commit the heat branchaifoundry1, repositoryPartly done aifoundry1's host programs rebuilt with the g3log fix; the PCIe probe's lock fixed (D24); the heat branch merged and recorded. Left: the card-0 guard's lock, less polling and fail-closed dry runs, which wait for the owner's decision on the heat-placement lock that freezes lib.sh and queue.shC8, C11, C13, C17, C22, C26, H26, D17, D24
U14Only with the owner's OK: our cron and linger, the add-lab-user.sh wording, the enercat_v2 rebuild, this page's visibilityvariousOwner's call Our cron and linger (on aifoundry2 our crontab now also runs a lab dashboard every 10 minutes and a session watchdog every minute), the add-lab-user.sh wording, the enercat_v2 rebuild. This page's visibility was settled on 26 Sep (public)H26
U15Reset aifoundry2's hung card (section 2.9)aifoundry2Done 28 Sep 08:32:45, with the owner's approvalC14, C27
U16et-reset: a checked card reset (section 2.9)all threeDone Installed on all three: aifoundry1 and aifoundry3 by the owner on 30 Sep, aifoundry2 at 15:36 that dayC6, C14
U17Scrub aifoundry1's pool, once the owners have dealt with their files (section 2.9)aifoundry1Waits for the owner's decision Not run: no scrub since 13 Sep (14:40:59 on 30 Sep); the automatic one is due on 11 OctH4
U18Run et-lab-health every day (section 2.9)all threeDone Rev 3 and the daily timer on all three: aifoundry1 and aifoundry3 by the owner on 30 Sep, aifoundry2 at 15:36 that dayC1, C3, C15, C16, C21, H3, H5
U19Reboot aifoundry1 and aifoundry2 into 7.0.0-34 (section 2.9)aifoundry1, aifoundry2Done Done by the power cycle of 30 Sep at about 15:07, after C28: all three run 7.0.0-34C2, H6
U20Quiet the desktop services (section 2.9)all threePartly done Done on all three (aifoundry1 28 Sep; aifoundry2 and aifoundry3 by the owner on 30 Sep at 14:37 and 14:39); the failed unit on aifoundry2 cleared with the reboot; apport's crash report there is left (root)C19, H33
U21After the campaign: aifoundry2's tmux, and the timers (section 2.9)aifoundry2; all threeWaits for the owner's decision Command and risk: section 2.9H25, H26
U22Turn off password logins in sshd, and review sudo (section 2.9)aifoundry1, aifoundry2Partly done Password logins off on aifoundry1 (28 Sep) and aifoundry2 (the owner, 30 Sep 14:37); the sudo review not runH19
U23OpenSSH on aifoundry3, key-only, for when Tailscale fails (section 2.9)aifoundry3Waits for the owner's decision Command and risk: section 2.9H24
U24Peer names in /etc/hosts on every lab machine (section 2.9)all threeDone Done on all three; aifoundry3's wrong line corrected at 15:36 on 30 SepH23
U25Test card 0's link at Gen3 before the visit (section 2.9)aifoundry1 card 0Ran; took aifoundry1 down Run by the owner on 30 Sep at 14:40–14:41: the retrain to 8 GT/s took the whole host down (C28). Not to be repeatedC15, C28
U26One boot with IOMMU passthrough, to test the copy rates (section 2.9)aifoundry2No longer needed E55 (29 Sep) refuted the IOMMU as the cause of the halving (D19)D19
U27Headless boot on all three (section 2.9)all threeDone Set on all three: aifoundry1 28 Sep; aifoundry2 and aifoundry3 by the owner on 30 Sep (14:37, 14:39); in effect on all three since their boots of 30 Sep and 2 OctH6, H8, H33
U28Install what was staged on 2 October: the corrected banners, the current et-lab-start, the root-login notice and et-opens (section 2.9)all threeWaits for the owner Staged and checked on 2 Oct; not installed on 4 Oct (H37)H37, C21, H28, D18

4.5 Still pending, and for whom#

WhatWhereWhoRequests or stepsProblems
Find why aifoundry2's Master Minion hung at 02:50 on 28 Sep, and why the per-card sysfs reset did not recover it (the management reset did, at 08:32); it has not recurred in the 52 launches of DV2's validationaifoundry2Firmware; the driverCF3C27, C14
Our root steps still open: installing the corrected banners and et-lab-start; removing apport's crash report on aifoundry2; the sudo review; OpenSSH on aifoundry3; tmux and the timers; the scrub, after the owners' filesall threeUs, on the owner's decision; changes on aifoundry3 go to its admin first (DI4); who may run et-reset is the lab's rule (PO2, CF10)U28, U20, U22, U23, U21, U17H37, C19, H19, H24, H25, H26, H4
Tell aifoundry3's admin about the 30 Sep changes there, the clock guard's drop-in and its boot markeraifoundry3RomanDI4C6, C2
Delete or move user data (no longer urgent: 116 GB free since 30 Sep); restore or delete the four corrupt files (the automatic scrub is due on 11 Oct); a backupaifoundry1Accounts rehan and roman; the owner of another home; the labMO1, MO2, PO4H3, H4
aifoundry2's card cooling, first (its card is off the bus); a second DIMM for aifoundry3. Card 0's fan was replaced on 2 Octaifoundry1, aifoundry2, aifoundry3On siteSH5, SH7H28, H29
Thermal protection, one firmware release (with the voltage-set fix), and whether card 1's DVFS is off on purposeall four cardsRoman; firmware; the lab (one short cool-start run on card 1)CF1, PO1, CF8C22, C4, C24, C12
Root between machines and the demo's open port; who keeps sudo and adm; Tailscale's check modeall threeRoman; the tailnet admin; aifoundry3's adminAS1, AS2, AS3H2, H13, H19, H1
The power cuts, a UPS, Ethernet, one console visit (MOK, SSD firmware, BIOS), an IP-KVMall threeOn siteSH2, SH3, SH4, SH6H7, C20, H20, H11, H21
The card lock as the rule; the CI runners and aifoundry3's demoall threeRoman; the CI owners; aifoundry3's adminPO2, MO3H31, H12, H13, H27
aifoundry3's clock pin, and its guard after a card resetaifoundry3aifoundry3's admin; firmwareCF2, DI4C23, C5, C6
The log-level race in the runtime; one reference /opt/et with a gp-sdkall three; et-platformNekko (runtime); the labRT1, RT2C17, C18, D10
Test Docker 29 and containerd 2 before unholding them; a standing maintenance windowaifoundry1The CI owners; RomanPO3H18
Adopt the public onboarding page (the banners link it since 30 Sep) and link it from #community-lab, and the per-card sheetall threeRomanDI1, DI2C10, H1, C4, C24
Publish the accounts page correctionspacesheepThe owner's OKU7H19, D18
Our card-0 guard's lock, less polling and fail-closed dry runs, and the lib.sh comment (after the owner's decision on the heat-placement lock); our live monitor: the card lock around each read, no per-second sudo, the libDM crashrepository; all threeUsU13H35, C26, D17, D23
Our cron and linger, the add-lab-user.sh wording, the enercat_v2 rebuildvariousThe ownerU14H26, D23, C17, H2

4.6 Checks after each reboot#

From the 25 September fix plan (B4). Read on 4 October on all three without opening a card: everything holds, except that aifoundry2's card is off the bus while its /dev nodes remain (C32). So check each card's link before the last line, which opens the card, and skip aifoundry2's card while it is out of service; on aifoundry1 that line opens both cards (C13), so run it only with both free. The et-who --check and et-lab-health lines were added on 28 September, when both were installed. Expected values after #.

uname -r                                                 # 7.0.0-34-generic
cat /sys/module/et_soc1/version /sys/module/et_soc1/srcversion   # 0.20.0 / 47D26A305A0428B29FB7FC4
ls -l /dev/et*                                           # crw-rw-rw- (aifoundry1: et0_* and et1_*)
systemctl is-system-running; systemctl list-jobs         # running / none (else: plymouth quit, H8)
chronyc -n tracking | grep -E 'Leap|System time'         # Normal
sysctl -n kernel.core_pattern                            # |/usr/lib/systemd/systemd-coredump ...
powerprofilesctl get; iw dev wlp8s0 get power_save       # performance / off
ls -l /run/lock/etsoc-*.lock; et-who --check; echo $?    # lock files recreated; 0 (no holder)
et-lab-health                                            # read-only; WARN lines name what is off
journalctl --list-boots | tail -2                        # the previous boot is kept
# aifoundry1: cat /sys/bus/pci/devices/0000:00:01.0/aer/correctable_ratelimit_interval_ms   # 60000
#             zpool status -x; grep RxErr /sys/bus/pci/devices/0000:00:01.0/aer_dev_correctable
# aifoundry3: cat /run/et-board-clock-guard.ok /proc/sys/kernel/random/boot_id   # same boot_id, 600 400 0
grep -H . /sys/bus/pci/drivers/ET/0000:*/current_link_speed     # 16.0 GT/s PCIe per card; Unknown: that card is off the bus (C32), skip it
# then, as an ordinary user with nobody on the card, per card N (not aifoundry2's while out of service):
/opt/et/bin/dev_mngt_service -m DM_CMD_GET_MODULE_FIRMWARE_REVISIONS -n N -u 5000

Measurements: the performance profile, the new kernel and the reboots change host-side timing (launch, sync and copy overheads), so host-side data from before about 16:15 on 25 September is not comparable with later data. Card-side results at a fixed operating point are unaffected by every change above.

5. How this was compiled#

Compiled by Claude (Anthropic) for Yaroslav Bulatov from the sources above; source of this page: docs/reports/2026-09-25-lab-problems.html in the owner's working copy, kept out of the public repository (it is gitignored) because it names other lab accounts. Related: the hub page on the limits of observability, whose ladder of improvements holds the firmware and documentation asks this list points to.

Appendix A. Draft: the onboarding page#

For the lab to adopt (request DI1): what every lab user needs before touching a card, from this page's problems. Written on 27 September by this page's author; it names no other accounts.

4 October: superseded by our public New user? Start here page, which newcomers follow; this draft is kept as of 27 September. Out of date in it: card 0 is back in service and aifoundry2's card is out of service (C21, H28); et-lab-health runs rev 3 every day; the short peer names now reach the tailnet (U24); and on aifoundry1 dev_mngt_service and et-powertop open both cards, so run them only with both free (C13).

Show the draft

Draft for the lab to adopt, 27 Sep 2026 (request DI1). Written by the lab-problems report's author from the report's problems and our notes; every section cites the problems it comes from. Once adopted, link it from the three login banners and keep it current. It assumes the 27 Sep versions of et-who (with --check) and et-lab-health; they are installed on all three hosts since 28 Sep (tools/lab/ in the repository).

Three x86_64 Ubuntu 24.04 machines, four ET-SoC-1 cards: aifoundry1 (two cards), aifoundry2 and aifoundry3 (one each). The cards are shared with other people, CI jobs, a demo service and long-running measurement queues. How the four cards differ is on the per-card sheet (DI2).

1. Getting in

  • Logins go through Tailscale SSH in check mode. ssh <you>@aifoundryN prints a https://login.tailscale.com/a/… link and waits. Open it in a browser signed in to the lab tailnet with your member account and approve it while ssh is still waiting. One approval lasts the tailnet's check period (12 hours by default); then you are asked again.
  • If the link answers "Error 404 … could not be located", sign out of login.tailscale.com, sign back in, and run ssh again for a fresh link.
  • A coding agent cannot approve the check. When its ssh prints the link, it has to hand the link to you and wait.
  • "The tailnet policy does not permit you" also means "no such account on that machine": ask for the account first.
  • From any lab machine, ssh to another lab machine by its short name resolves over the Wi-Fi LAN, not the tailnet: you reach password OpenSSH, or nothing on aifoundry3. Put each machine's full tailnet (MagicDNS) name in ~/.ssh/config as its HostName.
  • /opt/et/bin is on PATH in login shells on every machine.

Sources: H1, H23; docs/lab-access.md in github.com/yaroslavvb/et-soc1-prototyping.

2. Before you touch a card

  • Look first. et-who lists every process, of any user, that holds a card's device node or card lock. It never opens a card, so it is always safe. The login banner shows the same list. In a script, use et-who --check: it exits 0 if nothing is held, 1 if something is (your own lock included), 2 if the check failed. Do not parse the sentence plain et-who prints when nothing is held.
  • Each device node opens in one process at a time. A second opener gets "Device or resource busy", and the kernel log says Mgmt: Tried to open same device multiple times. So one sampler or monitor left running blocks everyone else's tools on that card.
  • Hold the card's lock for your whole run: flock -n /run/lock/etsoc-shire<N>.lock <command> (card N; -n fails at once if someone holds it). The lock is advisory: it protects you only from tools that also take it. The measurement queues, aifoundry3's clock guard and the CI configuration use the same path.
  • A card is busy when et-who shows a holder or its lock is held. Someone logged in is not the same as someone using a card. [This rule is the lab's decision: request PO2.] Claiming a machine on the lab's Discord channel ("using aifoundryN", then "released") is still the courtesy.
  • Keep each device-opening process under 10 s (timeout 10), and release the lock between separate tests. Long holds block everyone, and a hung card needs a reset that only the lab admin should do.
  • aifoundry1 has two cards. Programs built on aifoundry1 against its /opt/et open only card n when you set ET_DEVICES=<n> (they then see it as device 0). The stock /opt/et/bin/dev_mngt_service and et-powertop ignore it and open both cards' management nodes even with -n <N>: run them only when both cards are free. Card 0 overheats: no work on it until its cooling is checked.
  • Other users are part of the environment. CI runners (aifoundry1, aifoundry2) and aifoundry3's demo service can use a card at any time, and several accounts use aifoundry3's card. Check et-who before and after a run.

Sources: C11, C13, C21, H12, H13, H27, H30, H31.

3. Stopping tools, and the one-line drain

  • Stop tools and samplers with Ctrl-C or a plain kill, never kill -9. A management tool killed in the middle of a request leaves its reply queued on the card. The next program that opens the card dies of std::bad_function_call on that stale reply, and leaves its own behind, so retrying never recovers.
  • To clear it once (when the card is otherwise free; on aifoundry1 when both cards are free): /opt/et/bin/dev_mngt_service -m DM_CMD_GET_MODULE_POWER -n <N> -u 5000. It crashes on the stale reply and clears it; the next program starts normally.
  • Samplers sometimes fail to start right after a previous one stopped (about one start in three): retry until the output has a line.

Sources: C10, C19.

4. Long jobs

  • Start them detached so a dropped connection or an agent's command timeout does not kill them: ssh aifoundryN 'cd <dir> && setsid nohup <command> > <log> 2>&1 < /dev/null & echo $!'. Plain nohup keeps the ssh open. The machines are on Wi-Fi and drop now and then.
  • Stop a job by the PID you saved, not by pattern. pkill -f and pgrep -f over ssh match their own remote command line (the processes that run a remote command carry the whole command, including your pattern). When you must search, bracket the first letter: pgrep -af '[m]y_runner.sh'.
  • No secrets on the command line of an ssh command. Tailscale SSH logs every remote command line in the machine's journal, which admins can read.
  • Do not change code under a running job. Bash reads a script as it runs, so editing or scp-ing over it breaks the run; replace a file with mv onto the old name, and chmod +x after copying. Do not rebuild into a build directory a running job uses.

Sources: D14, D15, H11.

5. Files

  • /tmp is cleared at every boot. Keep work, logs and agents' scratch directories in your home directory.
  • Each machine has its own /home: files are not shared between machines.
  • aifoundry1's disk is one pool shared with the system, and it filled up once (September 2026). Keep one copy of each large model and delete what you no longer use.

Sources: H22, H3.

6. When something breaks

  • Read the kernel log first: dmesg | grep ET. Every user can read it. It says when a device open was refused (someone else holds the node), when the card reported an error event, and when a card was reset.
  • The driver's error counters are readable by everyone: cat /sys/bus/pci/devices/<BDF>/err_stats/ce_count (and uce_count). PmicCeEvent and ThermThrottleCeEvent mean the card hit its board-power limit or throttled for heat.
  • Core dumps are kept: coredumpctl list, then coredumpctl gdb <pid>.
  • et-lab-health is a read-only check of the machine and its cards (driver, links, error counters, disk, services); it warns about anything unusual and never opens a card.
  • A card looks wedged when every launch fails with KernelLaunchCmIfaceMulticastFailed or "Couldn't use the HPSQ. Perhaps the Master Minion is hanged?". Do not reset it yourself: tell the lab admin. On 28 Sep the management reset recovered such a card without a power cycle; the per-card sysfs reset did not (C27).

Sources: C14, C19, H15, H16.

7. Writing host programs

  • Register the runtime's log levels first in main, before anything creates the runtime. Otherwise about one launch in 100 on aifoundry3 crashes 1.08 s in (a race in libetrt's logging, present in every build):
    #include <g3log/loglevels.hpp>
    static void registerRuntimeLogLevels() {
      const LEVELS levels[] = {LEVELS(g3::kDebugValue - 100, "VERBOSE_HIGH"), LEVELS(g3::kDebugValue - 99, "VERBOSE_MID"),
                               LEVELS(g3::kDebugValue - 98, "VERBOSE_LOW")};
      for (const auto& l : levels) g3::only_change_at_initialization::addLogLevel(l, false);
    }
    int main(int argc, char** argv) { registerRuntimeLogLevels(); /* ... */ }

    A crash of an older binary (exit 139, 1.08 s in) is this race: repeat the launch.

  • Link the ET libraries with -Wl,-rpath,/opt/et/lib.
  • Each machine has a different build of the vendor runtime in /opt/et, so compare host-side timings (launch, synchronisation, copies) only within one machine. Card-side numbers compare across machines.

Sources: C17, C18.

8. Commands that change a card for every user

The device nodes are open to every account, so these work for anyone and silently change other people's runs. Do not use them without the lab's OK:

  • a card reset (DM_CMD_RESET_ETSOC; a software reset can hang a card);
  • the service processor's trace or log level;
  • a telemetry statistics reset (DM_CMD_SET_DM_STATS_RUN_CONTROL, ettelem sample --reset-ms): it restarts every reader's power averages;
  • temperature thresholds, the TDP level (DM_CMD_SET_MODULE_STATIC_TDP_LEVEL), the frequency (DM_CMD_SET_FREQUENCY), and active power management (turning it off disables the card's clock governor).

Also leave the firmware, the driver and /opt/et alone, and aifoundry3's clock guard, which is deliberate.

Sources: C12, C5.

9. Recording a measurement

  • et-lab-manifest prints the machine facts to save with every run: host, kernel, CPU and power profile, clock sync, driver version, each card's PCIe link, the runtime library hashes and aifoundry3's clock-guard marker. It does not open the card, so record the card's firmware from your own tool (dev_mngt_service -m DM_CMD_GET_MODULE_FIRMWARE_REVISIONS -n <N> -u 5000, while the card is free; on aifoundry1 while both cards are free).
  • Check the clock in every telemetry sample: the four cards run at different clocks for different reasons (per-card sheet).

Sources: C4, C18; docs/lab-access.md.

10. Coding agents

Agents follow the same rules. Before an agent works on a lab machine, tell it which machine and card it may use and whether any running job must be left alone. An agent that only writes code never touches a card: give it a dry-run mode, and check et-who afterwards. An agent cannot approve the Tailscale check, a spacesheep login or a push: it has to ask you.

Sources: D17, H1.

Appendix B. Draft: the four cards, one sheet each#

For the lab to adopt (request DI2): each card's firmware, DVFS state, TDP policy, idle state and quirks, and aifoundry3's pin. Written on 27 September by this page's author; it names no other accounts.

4 October: this sheet describes the cards on 27–28 September. Since then aifoundry1's card 0 is back in service (its fan was replaced on 2 Oct; it idles at 48–55 °C and 19–21 W with no error events), and its root port counts about 7 corrected errors an hour, not one a second (C15); aifoundry2's card is out of service, off the bus since 2 Oct (H28); and the “hottest sensor” values are peaks held since the card started (C30). The New user page's card table is current.

Show the draft

Draft for the lab to adopt, 27 Sep 2026, updated 28 Sep (aifoundry2) (request DI2). Written by the lab-problems report's author from the report's problems and the knowledge base in github.com/yaroslavvb/et-soc1-prototyping (docs/findings/14-card-behaviour.md first). Every line cites where it comes from. "From the source" marks what was read in the firmware source for that release and not tested on a card. Once adopted, update it after any reflash, card reset, policy change or service change (last section).

Today the four cards behave four different ways: aifoundry2 governs its clock, but only from a cool die; aifoundry3 is pinned at 600 MHz and its governor appears latched; aifoundry1's card 1 never changes clock; aifoundry1's card 0 idles at 300 MHz and overheats under load.

At a glance

aifoundry2aifoundry3aifoundry1 card 0aifoundry1 card 1
PCI address, nodes0000:02:00.0, /dev/et0_*0000:02:00.0, /dev/et0_*0000:01:00.0, /dev/et0_*0000:02:00.0, /dev/et1_*
Firmware release1.3.11.3.11.4.11.2.0
BL2 / PMIC / minion firmware0.20.0 / 1.5.0 / 0.23.00.20.0 / 1.5.0 / 0.23.00.21.2 / 1.6.1 / 0.24.00.18.0 / 1.3.0 / 0.22.0
Closest public sourceet-platform cafe03fc3^ (mid-May 2024)the same50310b06b (0.21.0; 0.21.1 and 0.21.2 are not public)da192816a (0.18.0, Mar 2024)
TDP: driver / firmware65 W / 65 W65 W / 0 W (set at every boot)65 W / 65 W65 W / 65 W
DVFSonpinned; governor appears latchedon, but acts only while a kernel runs (from the source)appears off
Boot clock (driver)600 MHz600 MHz600 MHz600 MHz
Clock in practice600 MHz; 700–800 only from a cool die600 MHz minion, 400 MHz NoC300 MHz idle, 600 MHz under load600 MHz always
Idle31–36 W at 73–80 °C (27 W cold)23.6 W at 50 °C; die idles at 55–57 °C18.6–18.8 W at 300 MHz; 62–63 °C33–35 W at 57–62 °C
Thermal step seenyes, from below about 68 °Cnone since 25 Sep10 throttle events and a drop to 300 MHz on 25 Sep, yet it reached 115–117 °Cnone, up to 88 °C
Use it forthe main cardswitching power over idle, never absolute wattsnothing: it overheatsanything, at a fixed 600 MHz

aifoundry2

  • Firmware: release 1.3.1: BL1 and BL2 0.20.0, PMIC firmware 1.5.0, master, worker and machine minion firmware 0.23.0, ASIC revision 597. The build is from about mid-May 2024; its closest public source is et-platform cafe03fc3^, not the December 2025 tree most readers read. (data/2026-09-25-claims-v3/firmware.md; C8)
  • TDP policy: 65 W, the same in the driver and in the firmware; temperature threshold 65 °C; power state managed_power. (E41 telemetry pass 1, raw/aifoundry2/tel/p1/gov/)
  • DVFS: on. Operating points 600 MHz at 516–518 mV, 700 MHz at 568 mV, 800 MHz at 618–620 mV. While a kernel runs, the governor steps down when the whole-degree mean die temperature is above 65 °C or board power is above the TDP, and up otherwise; the temperature test wins. Clock lifts were seen only below about 68 °C. (14-card-behaviour.md "The clock governor is thermal first"; E10, E29)
  • Boot clock: 600 MHz (the driver's configuration). (E41 driver.json)
  • In practice, 600 MHz. In this chassis the die never read below 65 °C in the version-3 campaign (336,070 samples, 25–26 Sep): blocks started at 67–99 °C, and 900 s of idle after heating cooled it only to 67–77 °C. It rested at 66–67 °C on 27 Sep. 700–800 MHz needs a cold start after a long idle (seen once, E10). (H28, C7)
  • Idle: 31–36 W at 73–80 °C; 27 W from cold. (14-card-behaviour.md)
  • Thermal protection: the governor's step-down only. From the source: it reads the mean of the shire sensors, not the hottest one; there is no die-temperature hard trip; the PMIC's 75 °C and 75 W alarms drop the clock to 300 MHz; and 1.3.1's safe state may not move the PLL (fixed upstream in e024210bc, after this build). In practice nothing tripped in a 26 Sep campaign pass at a 90–103 °C mean and up to 86.9 W (145 samples at 75 W or more): the clock stayed at 600 MHz. (C22)
  • Driver error counters (27 Sep 22:11): MinionCeEvent 23, SpCeEvent 5; no power or thermal events. (et-lab-health)
  • State on 28 Sep: its Master Minion hung at 02:50. The per-card sysfs reset at 06:39 did not recover it; the management reset (DM_CMD_RESET_ETSOC) at 08:32 did, and a test kernel ran at 08:33. (C27)
  • Quirks:
    • Some on-chip traffic starves the card's own meter: a telemetry sample takes 76–146 ms, or about 1 s, instead of 22 ms. Check took_ms. (D2; E27, E29, E31)
    • The service processor's management pass takes 133 ms with nothing polling. (E41)
    • 1.3.1's power state never reports low_power, and the maximum-temperature query returns 0. (C19, C22, from the source)
  • Host: the stock et-platform 353f20e build of /opt/et. A CI runner (as root) can use the card. (C18, H12)
  • Use it for: the main card. Preheat to about 76 °C and check mhz.minion in every sample.

aifoundry3

  • Firmware: identical to aifoundry2's: 1.3.1 (BL2 0.20.0, PMIC 1.5.0, minion 0.23.0, ASIC revision 597). (firmware.md)
  • The pin, and why. et-board-clock-guard.service, a root boot service in place since 23 Jul 2026, sets the static TDP level to 0 W (DM_CMD_SET_MODULE_STATIC_TDP_LEVEL) and the clocks to 600 MHz minion and 400 MHz NoC (DM_CMD_SET_FREQUENCY) at every boot. Its stated reason: the card "becomes unreliable when firmware DVFS raises the minion clock above its 600 MHz minimum", so it is kept "at the supported 600/400 MHz point by setting a zero-watt software TDP ceiling". The zero is not flashed. The driver still reports 65 W; the firmware reports TDP 0 W, threshold 65 °C, power state max_power. The guard writes /run/et-board-clock-guard.ok (boot id, 600, 400, 0); on 27 Sep it matched the current boot. (C5; data/2026-09-25-aifoundry1/facts.md; E41 raw/aifoundry3/)
  • DVFS: pinned, and the governor appears latched. At a 0 W TDP its power-down loop cannot finish at the 600 MHz floor, so after the first 65 °C crossing it makes no thermal step until the service processor restarts (from the source). Consistent with every trace since 25 Sep: no throttle or idle event where at least 5 were predicted (E41), and a silent trace probe on 27 Sep. Nothing slows a hot die: runs reach 88–90 °C. (C23, C5)
  • After a card reset the card runs its own DVFS until the guard runs again, and the guard trusts its boot marker instead of re-checking the card. (C6)
  • Boot clock: 600 MHz (driver). In practice: 600 MHz minion, 400 MHz NoC, never seen higher.
  • Idle: 23.6 W at 50 °C; the die idles at 55–57 °C since the host changes of 25 Sep (53–54 °C before). 150 back-to-back 2 s matmul launches took it from 55 to 88 °C. (14-card-behaviour.md; V3-IDLE)
  • Driver error counters (27 Sep): MinionCeEvent 1 (a kernel-launch error at 04:42, during another account's session while none of our work ran; the card recovered without a reset), everything else 0. (H31)
  • Quirks:
    • The management pass takes 224 ms against aifoundry2's 133 ms on the same firmware (possibly the stuck power task). (E41, C23)
    • Host programs crash about once in 100 launches, 1.08 s in, unless they register the runtime's log levels first (the host's patched -O3 libetrt.so hits a race in every build). (C17)
    • The host copies memory at 9.2 GB/s against 17–21 GB/s on the other two, so staged copies to the card reach 5.2 GB/s. The cause is the host's memory: one 32 GB DDR4-2666 DIMM on a single channel (read 27 Sep 22:18), where the other two hosts run two channels; a matching second DIMM in channel B should fix it (request SH7). (H29, E50; data/2026-09-27-pcie/hosts.txt)
    • The demo web service can launch jobs on the card as root, without the card lock; several accounts use the card. (H13, H31)
  • Use it for: comparing switching power over idle, never absolute watts. Cap every run's die temperature yourself (we stop at 90 °C).

aifoundry1 card 0

  • Firmware: release 1.4.1: BL2 0.21.2, PMIC firmware 1.6.1, minion firmware 0.24.0. Its exact source is not public: the closest is et-platform 50310b06b (0.21.0, 25 Sep 2024). (14-card-behaviour.md; C8)
  • TDP policy: 65 W in the driver and the firmware. Its idle power state reads low_power, which on 0.21+ only means board power at or below 30 W. (14-card-behaviour.md; C4, C19)
  • DVFS: on, but from the source the 0.21.x governor acts only while a kernel runs, so a hot idle die gets no software response; and 0.21.0 keeps board power in 16 bits, so above 65.535 W it steps the clock up instead of down (fixed upstream in 478275330, Nov 2024; unknown in 0.21.2). (C21, C22)
  • Boot clock: 600 MHz (driver). Idle: 300 MHz at 398 mV, 18.6–18.8 W; 62–63 °C (27 Sep). Under load it runs at 600 MHz (26 W idle there); the first launch after a low-power idle can be slow. (14-card-behaviour.md)
  • It overheats. On 25 Sep about ten minutes of short test launches took the die to 98–102 °C; just after, it read 115–117 °C with nothing running and drew 66–71 W at 600 MHz, until the firmware dropped it to 300 MHz. The driver counted 18 board-power events (75.0–75.75 W against a 75 W limit) and 10 thermal-throttle events (16:39–17:46), none since. Card 1 beside it peaked near 71 °C under the same launches. (C21)
  • PCIe: its root port logs about one corrected receive error per second, at the same rate with the card idle, so the physical link is suspect; no uncorrectable errors. (C15)
  • Quirks: the stock dev_mngt_service and et-powertop open both cards' management nodes, so a stock tool for card 1 collides with anything holding card 0. (C13)
  • Use it for: nothing, until its fan, heat sink and airflow are checked and it is reseated (request SH1). It was excluded from the version-3 campaign (amendment A4).

aifoundry1 card 1

  • Firmware: release 1.2.0: BL1 and BL2 0.18.0, PMIC firmware 1.3.0, minion firmware 0.22.0. Closest source: et-platform da192816a ("Close development of version 0.18.0", 27 Mar 2024). (E41 raw/aifoundry1-c1/tel/p1/gov/fw.txt; C8)
  • TDP policy: 65 W in the driver and the firmware; threshold 65 °C; power state managed_power. (E41 raw/aifoundry1-c1/)
  • DVFS: appears off. It read 600 MHz in all 359,657 telemetry samples of the campaign (25–26 Sep), including 11,446 below 65 °C at 45–63 W, where 1.2.0 should step up, and readings up to 88 °C, where it should step down. Its trace probe on 27 Sep was silent. The simplest reading is that active power management is off on this card (set by a management command, or by an invalid operating-point table at boot); we asked the lab (request CF8). (C24, C7)
  • Boot clock: 600 MHz (driver). In practice: 600 MHz always, and no thermal step.
  • Idle: 33–35 W at 600 MHz and 57–62 °C; 55 °C at rest on 27 Sep. (14-card-behaviour.md; C24)
  • Driver error counters (27 Sep): all zero. (C21)
  • Quirks:
    • Select it with ET_DEVICES=1 in programs built against aifoundry1's /opt/et (a May 2026 fork build); the stock tools open both cards. (C13, C18)
    • The management pass takes 135 ms. (E41)
  • Use it for: anything, at a fixed 600 MHz. Cap the die temperature yourself: nothing on the card will.

What all four share

  • The et_soc1 driver 0.20.0; the driver reports the same 65 W nameplate for every card, whatever the firmware uses. (C4)
  • A 600 MHz boot clock, and a PCIe 4.0 x8 link (16.0 GT/s) that trains fully on every card. (E41 driver.json; E50; C15)
  • On the link, every card and its root port run MaxPayload 256 B and MaxReadReq 128 B, with ASPM off; the host translates the card's DMA addresses (IOMMU, DMA-FQ). (read as root on 28 Sep; D19)
  • No thermal sensor the operating system can see: a card's temperature is readable only through its management node, which one process can hold at a time. (C25)
  • From the source: the governor's only thermal input is the whole-degree mean of the shire sensors, compared with > 65 and no dead band; the hottest sensor is not used; the maximum-temperature query returns 0; and there is no die-temperature hard trip, only the PMIC's alarms. (C22)

Keeping this sheet current

After a reflash, a card reset, or a change to a TDP, clock, power-management setting or boot service:

  1. Re-read each card's firmware release (dev_mngt_service -m DM_CMD_GET_MODULE_FIRMWARE_REVISIONS -n <N> -u 5000, with the card free) and the driver's view (et-lab-manifest; the repository's tools/etcfg).
  2. Re-read the firmware's TDP, threshold and power state, and whether active power management is on.
  3. Record a few minutes of clock and temperature at idle and under a short load.
  4. Update this sheet, the machine's login banner, and the date at the top.