| Cards, driver and firmware |
| C1 | aifoundry1's two cards were refused by every ET tool: the driver module had an empty version stringHolds through the 30 Sep and 2 Oct boots: all three hosts load the driver with its version string (0.20.0, the same srcversion), and aifoundry1's card nodes came up at the 2 Oct boot. What is left is the lab's: package the driver so it cannot recur, and a clearer error upstream. | aifoundry1 | blocks work | Fixed 25 Sep | Nekko team CF4, CF5; us U18 |
| C2 | Nothing reliable loaded the ET driver at boot on aifoundry1 and aifoundry3Tested at boot on all three: after the power cycle of 30 Sep at about 15:07 each host loaded et_soc1 0.20.0 by itself and created every card's nodes, aifoundry1 for the first time since its fix of 25 Sep. | aifoundry1, aifoundry3 | corrupts results | Fixed 30 Sep | Nekko team CF4, DI4; us U19 |
| C3 | Stale and foreign ET driver builds sat in DKMS, a risk at every kernel updateHolds. One install recipe (the package) is still the lab's. | all three | wastes time latent | Fixed 25 Sep | Nekko team CF4; us U18 |
| C4 | Four cards run three firmware releases and idle at different operating points; the driver reports the same nameplate for allNo sign of a reflash: /opt/et and the driver are unchanged (the firmware query was not run, since it opens the card). The dates of the /dev/et* nodes, which earlier checks used, are no evidence either way: they change at snap refreshes (aifoundry3 at 07:17 and aifoundry2 at 22:43 on 2 Oct, with nothing but snap and AppArmor in the kernel log). The kernel log is the check. | all four cards | corrupts results | Open updated | Nekko team PO1, DI2, DI3 |
| C5 | aifoundry3 is held at 600 MHz by a root boot service, not by its firmware, and the host cannot see itSame pin, new consequence: aifoundry3 has no working thermal governor (C23). | aifoundry3 | corrupts results | Open updated workaround | Nekko team CF2, DI2 |
| C6 | After a card reset, aifoundry3's clock guard trusts its boot marker and does not re-check the cardUnchanged: the guard script is as it was on 23 Jul, and aifoundry3's card has not been reset since 25 Sep 16:22. et-reset is installed on aifoundry3 since 30 Sep 14:39 (U16); it never touches the guard and prints a reminder to tell aifoundry3's admin instead, so after a reset the 600 MHz pin still has to be restored by that admin (CF2, DI4). | aifoundry3 | corrupts results latent | Open | Nekko team CF2, CF3, DI4; us U16 |
| C7 | On cards with a free governor the clock follows the die temperatureaifoundry2's governor follows the 34-sensor mean with no dead band (DV2's validation, 28–29 Sep), but its die idles above 65 °C, so it rarely acts (H28); card 1 never steps (C24), aifoundry3 is pinned (C23), and card 0 was not measured. | measured on aifoundry2 only (from a cool die); aifoundry1 card 0 unmeasured, but it throttled 10 times on 25 Sep | corrupts results | Open updated workaround | Nekko team CF3, CF8, SH5 |
| C8 | The firmware on the cards is a mid-2024 build, but the source everyone reads is from December 2025We mapped all four cards to their source, and the mapping, with each build's governor, has been in the repository since 28 Sep (docs/findings/14-card-behaviour.md); 1.4.1's exact source is still not public (CF9). | all four cards | corrupts results | Open updated workaround | Nekko team CF9, PO1 |
| C9 | Users cannot run changed firmware: the images must be signedUnchanged. The firmware changes the measurements need are rungs 11–16 of the hub's ladder. | all four cards | blocks work | Open | Nekko team CF12 |
| C10 | A management tool killed mid-request poisons the card's management queue, and samplers often fail to startNo incident since 25 Sep 17:00, and no vendor fix; the banners keep the never-kill-9 rule and the drain line. | all three | blocks work | Open workaround | Nekko team CF5, DI1 |
| C11 | Nobody could see who held a card node, which only one process can openHolds. Its follow-up H30 was fixed on 28 Sep (et-who --check, installed on all three); C26 remains. Upstream: name the holder, and a read-only telemetry path. | all three | wastes time | Fixed 25 Sep | Nekko team CF5, CF6, PO2; us U13 |
| C12 | Any user can change a shared card's global state: reset it or change its trace levelWider than reported: stats reset, thresholds, DVFS-off, TDP and frequency are open to every user too, and, from the firmware source (read for E54 on 28 Sep, untested), so is a rail's voltage: on card 1's 1.2.0 each NoC voltage set rewrites the flash sector the boot voltage comes from, and on the two 1.3.1 cards a failed set loops until the watchdog resets the card (CF10, PO1). Since 28 Sep the banners on all three hosts list the other commands (U5), not yet the voltage set. | all three | corrupts results | Open updated | Nekko team CF10 |
| C13 | On the two-card host the stock ET tools open both cards; picking one needs aifoundry1's forked runtimeNew cost, from our own tools: a stock tool for card 1 collides with anything on card 0. | aifoundry1 | blocks work | Open workaround | Nekko team CF7, DI4; us U13 |
| C14 | A hung card is recovered by power-cycling the whole host, although a per-card reset exists and worked on an idle cardNo hung card since 28 Sep (C27): the restored card ran DV2's 52 validation launches (28–29 Sep) without one. A checked reset command, et-reset, is installed on all three hosts since 30 Sep (U16; aifoundry2's at 15:36), for root and the sudo group. aifoundry2's card is now off the bus for another reason, its cooling (H28, C32): no reset can help it, and none should be tried before the cooling is fixed. Which reset is the supported recovery is still CF3's question, and who may run it the lab's (PO2). | all three | blocks work | Open updated | Nekko team CF3, PO3; us U16 |
| C15 | aifoundry1 card 0's PCIe link logs about one corrected error per secondThe flood has not come back, but the count is not zero. Since aifoundry1's boot at 12:59 on 2 Oct (after the fan visit) card 0's root port has counted 336 corrected receiver errors in 46.7 hours, about 7 an hour, against about 3,600 an hour before 30 Sep; card 0 itself and card 1's port count none, and both links run at 16 GT/s x8. The kernel logs them at 2–17 lines an hour. Our dashboard shows this port's rate as 0 an hour, which is wrong (ours to fix). Watch it; on site only if the rate grows. | aifoundry1 card 0 | wastes time | Partly fixed updated | us watch it (U18), fix the dashboard's rate; on site only if it grows |
| C16 | The PCIe error flood filled aifoundry1's logs on a nearly full diskHolds. On 28 Sep we cleared the systemd reload warning (daemon-reload: 0 units need one) and compressed the flood-era kern.log.1 and syslog.1 (287 and 299 MB to 11.7 and 13.6 MB, content checked). The size cap stays off on purpose. | aifoundry1 | wastes time | Fixed 25 Sep | Us U18 |
| C17 | Host programs on aifoundry3 crash about once in 100 launches, 1.08 s after startFixed in our programs on aifoundry2 and aifoundry3 (26 Sep: 641 processes, no crash); aifoundry1's were rebuilt on 28 Sep (six at 07:52, build/sparsity at 20:56), all but the campaign's enercat_v2; the runtime is unchanged, and another account now runs programs on aifoundry3 (H31). | aifoundry3 | wastes time | Partly fixed workaround | Nekko team RT1, DI1; us U13 |
| C18 | Each host runs a different build of the vendor runtime in /opt/etUnchanged: the three runtime builds still differ; E50 now shows that host-side timings follow each build. | all three | corrupts results | Open | Nekko team RT2 |
| C19 | The vendor tools mislead when something is wrongTwo more misleading readouts found (the maximum temperature reads 0; “low_power” is only a power threshold). apport's coredump hook, which wrote duplicate crash reports of our programs, is off on all three hosts since 30 Sep (aifoundry1 since 28 Sep; U20). On aifoundry2 the hook's failed unit cleared with the reboot, and the host reads running (4 Oct); only apport's own crash report of 28 Sep is left in /var/crash, for root to remove (U20). | all three | wastes time | Open updated | Nekko team CF5, CF1; us U20 |
| C20 | The driver's signing key is not enrolled: turning Secure Boot on would make the cards disappearNew detail: the firmware is in Setup Mode, so a BIOS reset that restores the default keys could turn Secure Boot on and hide the cards. | all three | blocks work latent | Open | Nekko team SH3 |
| C21 | aifoundry1's card 0 overheats under load: 98–102 °C in short test runs, 115–117 °C just afterFixed on 2 October: on site, card 0's fan was found broken and replaced (host up at 12:59). Our acceptance test that afternoon: 49 °C idle (the peak since the restart 53 °C) against 65 °C before, and 8 minutes of sgemm bursts (8 s under the lock, 2.5 s gaps) held it at 52–53 °C (peak 56 °C), cooler than card 1 under the same test (59–60 °C, peak 63 °C). The owner put it back in service, and the new-user brief now offers it (SH1). 4 Oct: it has idled at a 48–55 °C mean for two days (peak 57 °C), at 19–21 W, with no error events. The “123 °C hot spot” quoted before was a peak held since September (C30). The installed login banner still warns against card 0 until the corrected one is installed as root (U28). | aifoundry1 card 0 | blocks work | Fixed 2 Oct updated | us install the corrected banner (U28, root) |
| C22 | No card has a die-temperature hard trip or any hot-spot protection, and the lab's firmware releases carry known thermal and power bugsFound on 27 Sep in the firmware source of all three releases, and seen in our 26 Sep data (a pass at a 90–103 °C mean, up to 86.9 W, with the clock never leaving 600 MHz). 2 October: an idle card heated to 138 °C (peak sensor reading 144 °C) at 134 W, and nothing on the card acted; it dropped off the PCIe bus at 12:02 (H28). At idle a clock cut cannot help: the power is leakage, which grows with temperature (about doubling every 23 °C above a 13.6 W floor). Only a hard trip that lowers the minion voltage or cuts the minion rails, at about 105 °C, would stop a runaway; until the firmware has one (CF1), a host whose card loses its cooling has no protection. | all four cards | blocks work | New workaround | Nekko team CF1, PO1; us U1, U13 |
| C23 | aifoundry3's 0 W TDP pin also stops its governor, so the card makes no thermal stepInferred from the source; every service-processor trace on aifoundry3 since 25 Sep is empty, as predicted. | aifoundry3 | corrupts results | New workaround | Nekko team CF2, DI2 |
| C24 | aifoundry1 card 1's clock never moves: its DVFS appears to be off, so it has no thermal step either600 MHz in all 359,657 campaign samples, cool or hot, and again in all 10,033 samples of our 29 Sep runs; its trace probe was silent (27 Sep 20:24). 4 Oct: no command reads the active-power-management flag (the management API has only the set, which changes the card for everyone), and 1.2.0's closest public source (da192816a) turns it on at every service-processor boot with no VMIN check, so neither hypothesis below explains a card that never steps. Card 1's service processor restarted at the 2 Oct boot; a busy run on it that afternoon (13:07–13:15, 54–60 °C) logged no clock. One short cool-start run that logs the clock settles whether it steps now (CF8). | aifoundry1 card 1 | corrupts results | New workaround | Nekko team CF8, PO1, DI2 |
| C25 | The operating system cannot see a card's temperature, or any chassis fan: only the single-opener management node reports itChecked on 27 Sep and again on 4 Oct on all three hosts: no ET entry in hwmon, no fan readings. Since 2 Oct the lab dashboard's Live section and the History page show each card's mean temperature once a second, read by our live monitor through the management node (H35). | all three | wastes time | New workaround | Nekko team CF6 |
| C26 | Our own tools opened aifoundry1 card 0's management node without card 0's lockThe heat guards have not run since 28 Sep. But since 2 Oct our live monitor opens every card's management node, card 0's included, about once a second without the card's lock, while et-who --check shows no holder (H35). | aifoundry1 card 0 | wastes time | New | Us H35, U13; Nekko team CF6 |
| C27 | aifoundry2's Master Minion hung on 28 Sep; the sysfs reset did not recover it, the management reset didRecovered on 28 Sep: the per-card sysfs reset at 06:39 re-attached the card but left the Master Minion hung (launches at 06:41–06:47 still failed); the management reset at 08:32:45 recovered it, and a test kernel ran 3 launches at 08:33. It has not hung since: DV2's validation ran 52 launches on it on 28–29 Sep, all returning 0, one of them meeting the same clock step down as the hung launch. The cause, and which reset is the supported recovery, are CF3's questions. | aifoundry2 | blocks work | Fixed 28 Sep | Nekko team CF3 |
| C28 | Retraining aifoundry1 card 0's PCIe link to 8 GT/s took the whole host down on 30 Sep; it needs a power cycle on siteOur Gen3 test of card 0's link (U25) retrained its root port to 8 GT/s at 14:41 on 30 Sep, and the whole host froze at once: the journal's last entry is at 14:41:17, with no kernel error, machine check or panic record. Roman power-cycled the lab at about 15:07, and it came back with both cards working (the incident and its lesson). | aifoundry1 | blocks work | Fixed 30 Sep | us U25 (never repeat) |
| C29 | A race in the ET driver: reading a card's message counters while the card is reset, or while the driver loads, can return garbage or crash the readerFound on 30 Sep by reading the driver's source while reviewing our new usage logger; never seen on a card, and not tried. The driver shows these counters before it builds the queue tables they read, and frees the tables before it removes the counters. Our usage logger already ignores an impossible jump. 4 October: filed upstream, publicly, as et-platform issue #136. The chips are end-of-life, so the owner chose a public report. et-platform's head is still 836a4ab of 17 Jul, and the driver on all three hosts is unchanged (0.20.0, the same srcversion). A patch is ready for the maintainers (CF13). | all three | corrupts results latent | New workaround | Nekko team CF13 |
| C30 | The cards report one die temperature: the 35 sensors' own readings stay inside the firmware, and the lab's “hottest sensor” is a peak held since the card startedFound on 4 Oct. The host gets the integer mean of 34 minion-shire sensors plus peak-holds kept since the card's service processor started. Our dashboard shows that peak as the “hottest sensor”: aifoundry1's card 1 has shown 63 °C since 2 Oct at a 53–61 °C mean, and card 0's “123 °C hot spot” at a 65 °C idle was a peak left from September (C21). No command, trace or debug path gives the sensors one by one. | all four cards | corrupts results | New workaround | Nekko team CF14; us the dashboard's labels |
| C31 | A rejected read-only management query logs “Critical, SP Runtime Error”, although the card is fineFound on 4 Oct, reading the service processor's source against the kernel log. Every line the SP logs at error level counts as a runtime error, and with the threshold at 0 each one reaches the host as a Critical event. aifoundry2's first five such events (20, 22 and 25 Sep) each fell in the same second as a batch of our read-only queries, one of which the firmware rejected; the card ran kernels after each. Only the sixth, on 28 Sep, came with a hang (C27). | all four cards | wastes time | New | Nekko team CF3 |
| C32 | A card that falls off the PCIe bus still looks present, and a warm reboot does not bring it backaifoundry2's card dropped off the bus at 09:42 on 1 Oct and at 07:42 and 12:02 on 2 Oct, and has been off since. Each time its /dev nodes stayed, the driver stayed bound, and et-who called the card free; the kernel log showed only queue errors (SQ[0] sync: head mismatched, head_remote: -1), never “card lost”. Only sysfs tells: the card's link speed reads Unknown and its width 63 (still on 4 Oct). A plain reboot at 10:47 on 2 Oct left the slot empty; reboots that cut the slot's power brought the card back. | aifoundry2; any card | wastes time | New workaround | Nekko team CF5; us H34 |
| C33 | Nothing reads the cards' DRAM temperature, and DRAM refresh never speeds up on a hot cardFound in the source and in one measurement. The memory set-up leaves the controllers' temperature derating commented out (“not needed for bring-up”), nothing reads the LPDDR4X's own temperature register (MR4), and the memory shires have no sensor. In E53 (28 Sep) refresh stayed at its programmed rate at die means of 52–75 °C. LPDDR4X needs faster refresh above 85 °C, and the dies have run at 90–138 °C. No wrong result was seen up to an 81 °C mean; nothing hotter was checked. | all four cards | corrupts results latent | New | Nekko team CF15 |
| The host machines and access |
| H1 | Tailscale SSH asks for a browser check, and the check link can 404Unchanged. The check-mode steps and the 404 fix are in our repository's docs/lab-access.md and in Appendix A, but not yet on the public New user page that newcomers now follow (ours to add, AS3). | all three | blocks work | Open workaround | Nekko team AS3, DI1 |
| H2 | Any account on one lab machine can log in as root on the othersUnchanged on 1 Oct. Since 2 Oct the lab's own onboarding uses the shared root login on purpose: each newcomer logs in as root once to create their own account (the lab's message to newcomers, and our New user page), so any narrowing must keep a way to create accounts (AS1, AS5). | all three | security | Open | Nekko team AS1, AS5 |
| H3 | aifoundry1's disk is full of user data: 95% of a single 452 GB poolFixed on 30 Sep at 22:25 PDT, at the owner's word: we deleted a departed user's public model checkpoints, the 27 files of 1 GB or more (118.1 GB of public models such as Qwen, Llama, Gemma, SmolVLM, RWKV, LFM and TinyLlama, which can be downloaded again). The account no longer existed, and nothing used the files (no process had them open, and no system setting or scheduled job named them); the smaller checkpoints, all code and a 10 GB compiled bundle were kept. /home went from 99% used (7.3 GB free) to 72% (116 GB free), and the pool from 95% to 71% (130 GB free of 452 GB). The other owners' data is unchanged (MO1); with no quotas the pool can fill again (PO4). Still so on 4 Oct: 115 GB free on /home (72%), and the pool 71% full with 128 GB free. | aifoundry1 | blocks work | Fixed 30 Sep updated | Nekko team PO4 (so it does not fill again), MO1 (the rest, no longer urgent); us U18 |
| H4 | aifoundry1: ZFS permanent errors in four files, on a single disk with no backupsUnchanged: the 13 Sep scrub's 42 errors, no scrub since, no snapshot and no backup. The automatic scrub runs at 00:24 on Sunday 11 Oct. aifoundry1 was most likely switched off without a clean shutdown at about 12:55 on 2 Oct for the fan work (its journal stops mid-stream), another hard stop for this single disk. | aifoundry1 | corrupts results | Open | Nekko team MO2, PO4; us U17 |
| H5 | aifoundry1's 24 September kernel update stopped half-wayHolds: dpkg --audit is empty, and unattended upgrades ran cleanly on 26 and 27 Sep. | aifoundry1 | blocks work | Fixed 25 Sep | Us U18 |
| H6 | The reboot into kernel 7.0.0-34 has been pending since 24 Sep; done only on aifoundry3All three hosts run 7.0.0-34: aifoundry3 since 25 Sep, and aifoundry1 and aifoundry2 since Roman power-cycled the lab at about 15:07 on 30 Sep, after our link test hung aifoundry1 (C28). Section 4.6's checks passed on all three at 15:18–15:20. | aifoundry1, aifoundry2 | wastes time | Fixed 30 Sep | Nekko team PO3 (a maintenance window for the next one) |
| H7 | Most unclean resets hit all three machines at once: their power is cut togetherNo unplanned cut since 18 Sep. All three restarted together, without shutdown records, at 15:07 on 30 Sep: Roman's deliberate power cycle after C28. aifoundry1 was most likely switched off without a clean shutdown at about 12:55 on 2 Oct for the fan work. A single machine should be powered off by itself, after sudo poweroff (SH2). | all three | corrupts results | Open | Nekko team SH2 |
| H8 | The boot never finishes: the splash screen waits forever on headless machinesFixed: all three hosts boot to the text target (U27), and after the power cycle of 30 Sep at about 15:07 all three reported running with no queued jobs (15:18–15:20). | all three | cosmetic | Fixed 30 Sep workaround | nothing left |
| H9 | aifoundry1's journal kept only about a dayaifoundry1's journal is at its 1 GB cap again and has grown by about 170–185 MB a day since 2 Oct, mostly sudo lines from our own live monitor, one a second (H35). It still keeps about ten days, about five at this rate. aifoundry2's is at 1.14 GB of its 2 GB cap. | aifoundry1 | wastes time | Fixed 25 Sep | us H35 (stop the per-second sudo); aifoundry2's cap (root) |
| H10 | Time came from one NTP server over Wi-FiHolds: chrony is in sync with 8 sources on all three (offsets 0.09–0.33 ms). | all three | corrupts results | Fixed 25 Sep | nothing left |
| H11 | No Ethernet: all three machines run on Wi-FiThe mitigation holds (Wi-Fi power saving off); still no cable, and aifoundry2's roaming got worse: it switched access points 45 times on 30 Sep by 14:47 (3–25 a day on 26–29 Sep). | all three | wastes time | Partly fixed | Nekko team SH4 |
| H12 | CI runners run as root and can take a card at any timeUnchanged: both runners run as root, enabled and idle (no job since 15 Jun and 24 Jul), still polling GitHub. | aifoundry1, aifoundry2 | corrupts results | Open | Nekko team MO3, PO2 |
| H13 | aifoundry3's demo web app can launch jobs on the card as root, outside any lockUnchanged: the demo and the chatbot still run, with no requests since 25 Sep; disabling them no longer affects the driver (C2). | aifoundry3 | corrupts results | Open | Nekko team MO3, AS1 |
| H14 | numpy, venv, scipy and build packages were missing; aifoundry1 could not make a venv at allHolds: numpy 1.26.4, scipy and venv on all three. | aifoundry1, aifoundry3 | blocks work | Fixed 25 Sep | nothing left |
| H15 | Users could not read the kernel log, where the driver explains refused opens and card errorsHolds. (Users still cannot read the system journal: by design.) | all three | wastes time | Fixed 25 Sep | Nekko team DI1 |
| H16 | Crashing programs left no core dumpsProven in use. Side effect: apport's coredump hook also wrote .crash copies (C19) until it was switched off, on aifoundry1 on 28 Sep and on aifoundry2 and aifoundry3 on 30 Sep (U20); aifoundry2's failed hook unit cleared with its reboot; only apport's crash report of 28 Sep is left there (C19). | all three | wastes time | Fixed 25 Sep | Nekko team DI1 |
| H17 | Host CPUs ran in power-saving mode, so host-side timing jitteredHolds: the performance profile on all three. | all three | corrupts results | Fixed 25 Sep | nothing left |
| H18 | Months of pending updates, and a plain apt upgrade over Tailscale SSH kills itselfHolds; recurs by design without a maintenance window. Docker 29 / containerd 2 are held on aifoundry1 only, for the CI owners (aifoundry2 and aifoundry3 already run Docker 29.1.3 and containerd 2.2.1). On 4 Oct no security update was pending on any host; 28, 21 and 16 others were, among them kernel 7.0.0-38 (PO3). | all three | wastes time | Fixed 25 Sep | Nekko team PO3 |
| H19 | Sudo, passwords and OpenSSH do not match the documented policySudo grew on 2 Oct: 7 members on aifoundry1, 18 on aifoundry2 and 3 on aifoundry3 (4 Oct). One new account got sudo on all three hosts that day, and a new file appeared in aifoundry2's /etc/sudoers.d; the New user page promises accounts with no sudo. Password logins stay off on aifoundry1 and aifoundry2. Our public accounts runbook still says to give sudo by setting a starting password (U7). | aifoundry1, aifoundry2 | security | Partly fixed | Nekko team AS2; us U7, U22 |
| H20 | Host firmware: 2021 BIOS versions, and SSD firmware with a known health bugUnchanged: BIOS F5, F5 and F6, and SSD firmware 3B2QGXA7 on all three. Since the 30 Sep boot only aifoundry1 has the IRQ 9 storm (aifoundry2's interrupt rate now matches aifoundry3's, so the storm comes and goes), and all three log the same ACPI BIOS error at boot (\ADBG, AE_ALREADY_EXISTS). | all three | corrupts results latent | Open | Nekko team SH3 |
| H21 | No console or out-of-band access, and the boot menu was hiddenThe 5 s boot menu holds; still nobody at a console, and on 30 Sep it mattered: aifoundry1 went down and nothing could restart it remotely (C28, SH8). | all three | wastes time | Partly fixed updated | Nekko team SH6, SH8 |
| H22 | /tmp is wiped at boot, and coding agents keep their working files thereOur working files now live in the home directory (our rule since 30 Sep), so there is no /tmp copy to make, and U1 is closed. The rule itself is by design. | aifoundry2 (all three) | corrupts results | Partly fixed updated | Nekko team DI1 |
| H23 | On every lab machine, ssh to another lab machine goes over the LAN, not TailscaleFixed on all three: each machine reaches the other two at their tailnet addresses (U24); aifoundry3's wrong line was corrected at 15:36 on 30 Sep. | all three | wastes time | Fixed 30 Sep updated | nothing left |
| H24 | aifoundry3 has no sshd: Tailscale is the only way inUnchanged: no openssh-server on aifoundry3. Installing it key-only is ours now (U23, waiting for the owner's decision). | aifoundry3 | blocks work latent | Open | Us U23 |
| H25 | aifoundry2's tmux is a third-party snap that the Ubuntu package would breakUnchanged: the snap tmux stays held. Replacing it after the campaign is ours now (U21). | aifoundry2 | wastes time | Open workaround | Us U21 |
| H26 | Background load on the measurement hosts, some of it oursIt grew again on 2 Oct, and the largest part is ours: our live monitor runs on all three hosts and, every second, checks the holders with sudo and reads each card's temperature through its management node (H35); the dashboard (every 10 minutes) and the History page (every 5 minutes) run from aifoundry2's crontab with probes into the other two hosts, with a watchdog and a node watcher every minute. Pausing all of it during a campaign is ours (U14). | all three | cosmetic | Open | Us H35, U13, U14, U21 |
| H27 | An idle login blocks anyone who follows the "nobody else logged in" etiquetteIt recurs with the new users: on 4 Oct two newcomers' logins had been idle since 2 Oct (on aifoundry1 and aifoundry3), and two of ours on aifoundry1. Coding agents now run in tmux by design, so only the card lock can be the rule (PO2, H36). | aifoundry1, aifoundry3 | wastes time | Open workaround | Nekko team PO2; us close our idle sessions |
| H28 | aifoundry2's card never cools below the governor's 65 °C threshold, so its DVFS is almost never seenOut of service. The card has been off the PCIe bus since 12:02:01 on 2 Oct (its link reads Unknown, width 63), and the host has stayed on since 10:53 that day with the card in it. Its own temperature cannot be read; the host's drive, network-chip and CPU sensors have tracked aifoundry3's within about 2–3 °C since about 12:30 that day, so there is no sign that the card still heats. There is no workaround: only the cooling fix (SH5), then a power-cutting reboot to bring the card back (C32). | aifoundry2 | corrupts results | New | Nekko team SH5 (the next visit) |
| H29 | aifoundry3's host copies memory at about half the other hosts' rate: it runs on one memory channelCause found on 27 Sep at 22:18 and confirmed as root on 28 Sep (dmidecode): one 32 GB DDR4-2666 DIMM, in ChannelA-DIMM1, three slots empty; a second DIMM is on-site work. | aifoundry3 | corrupts results | New workaround | Nekko team SH7 |
| H30 | et-who prints a sentence when nobody holds a card and always exits 0, so a script took “free” for “held”Fixed on 28 Sep: et-who --check (0 free, 1 held, 2 failed) and the new idle sentence are installed on all three hosts; plain et-who still prints the same holder lines and exits 0. | all three | wastes time | Fixed 28 Sep | nothing left |
| H31 | Other accounts now work on aifoundry3, and one used its card; the report assumed only we didSeveral users at once is now normal: three people were onboarded on 2 Oct, and since then two other accounts have used aifoundry3's card and one aifoundry1's card 1 (the card-usage log). The newcomers took the card lock every time; one established account did not (PO2). | aifoundry3 | corrupts results | New workaround | Nekko team PO2, DI1 |
| H32 | aifoundry3's journal has reached its 2 GB cap and will start deleting the July boot historyFixed on 28 Sep, and at risk again: since 2 Oct the journal grows by about 140–225 MB a day, about ten times the rate of 28 Sep, mostly sudo lines from our own live monitor (H35). At 2.6 GB of the 4 GB cap it fills around 11–13 Oct and then deletes the July boots again (the export of 28 Sep survives). | aifoundry3 | wastes time | Fixed 28 Sep | us H35 |
| H33 | Desktop services on the headless hosts fill the error logThe desktop part holds on all three: bluetooth, cups-browsed and the updater are off, and all three have booted headless since. The error counts have not been re-read (that needs the adm group). The new noise is ours: our live monitor's core dumps and its per-second sudo lines (H35). | all three | cosmetic | Partly fixed workaround | Us H35; re-read the error counts (root or adm) |
| H34 | et-who --check covers the whole host and cannot see a card that is downet-who --check exits 1 if anyone holds any card node or lock on the host. On aifoundry1, where two people may now work at once, one per card, a script that gates on it waits for the other person's card. It lists holders only, so a card that fell off the bus keeps its /dev nodes and shows as free (C32). Each call runs sudo (H35). The banners and the onboarding brief say to read only your own card's lines and to check the link in sysfs. | aifoundry1, aifoundry2 | wastes time | New workaround | us et-who --card N with a link check (installing it needs root) |
| H35 | Our live monitor reads every card once a second without the card's lock, crashes now and then, and logs a sudo line every secondSince 16:05 on 2 Oct our live collector, a user service on all three hosts, reads each card's temperature once a second with a copy of ettelem, but only while et-who --check finds no holder on the host. Each read holds the single-opener management node for about 4 ms without the card's lock: about 177,000 opens of each of aifoundry1's cards by 4 Oct, and a tool started in that window gets “busy”. The reader has crashed in the vendor's libDM.so six times on aifoundry1 and twice on aifoundry3 (no reading lost), the failure C10 warns of. And each tick runs sudo, so aifoundry1 and aifoundry3 log about 3,600 sudo lines an hour, and their journals grow by 150–225 MB a day (H9, H32). | all three | wastes time | New | us take the card lock around each read, read the holders without sudo, report the crash (CF5) |
| H36 | The “is anyone else here” checks miss people, and nothing holds a card between runswho reads utmp, which has no entry for a Tailscale SSH command without a terminal: on 28 Sep it listed nobody on aifoundry3 while uptime counted three users. loginctl misses coding agents in tmux under linger: on 30 Sep aifoundry2 showed no sessions while two users had processes running, and every account made since 2 Oct has linger. The card lock lasts one run, so a newcomer setting up or between runs holds nothing. | all three | wastes time | New workaround | Nekko team PO2; us our repository docs |
| H37 | The root installs staged on 2 October were never made: two login banners contradict the cards' stateaifoundry1's login banner (30 Sep) still says card 0 overheats and must not be used, though its fan was replaced on 2 Oct and the dashboard and the New user page offer it. aifoundry2's does not say its card is out of service, and still invites newcomers to et-lab-start and the card lock. The installed et-lab-start (2 Oct 14:33, all three) hands the agent over at step 4 of the brief, which skips the simulator check. aifoundry1's card-usage logger skipped card 0 until 4 Oct; it logs it again. Corrected copies have been staged since 2 Oct and need one root session (U28). | all three | wastes time | New | us U28 (root) |
| Developing and measuring |
| D1 | The card's meters are coarse and filtered, and half of an idle card's power is on no meterNumbers replaced by E41, and the filter by E58 (29 Sep, pre-registered): a first-order average of 1.01–1.06 s on aifoundry3's rails and 1.08 s on card 1's minion and NoC rails, but 0.54 s on card 1's SRAM rail, so the cards' meters differ; aifoundry3's slower pass has a candidate cause (C23). | all four cards | corrupts results | Open updated workaround | Nekko team CF11, DI3 |
| D2 | Some on-chip traffic starves the card's own meter on aifoundry2Unchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep. | aifoundry2 | corrupts results | Open workaround | Nekko team CF11 |
| D3 | Power follows die temperature, heat carries over between runs, and the room's airflow changesaifoundry2's chassis keeps its card hot enough to hide its governor (H28). | all four cards | corrupts results | Open updated workaround | Nekko team SH5 |
| D4 | Common instructions trap in user mode: divide, square root, sine, 64-bit integer-to-float, double, the cycle CSRUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep. | all four cards | wastes time | Open workaround | Nekko team RT4, CF11, DI3 |
| D5 | Scratchpad addressing traps: offset 0 faulted once, and a global atomic through self ID 0x7F is a bus errorUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep. | all four cards | wastes time | Open workaround | Nekko team DI3, CF11 |
| D6 | The L1 data cache is not coherent and writes back whole 64 B linesUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep. | all four cards | corrupts results | Open workaround | Nekko team RT3 |
| D7 | One hot line stops a shire: hammering a global atomic stalls its home shire's memory path, with no errorUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep. | all four cards | blocks work | Open workaround | Nekko team RT3, DI3 |
| D8 | VPU register-file erratum 1.29: the compiler inserts no workaround and the simulator does not model itUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep. | all four cards | corrupts results | Open | Nekko team RT4, DI3, CD1 |
| D9 | sys_emu and silicon disagree in ways a new user does not expectUnchanged: sys_emu and the cards are as on 25 Sep. | sys_emu | wastes time | Open workaround | Nekko team RT6 |
| D10 | Current gp-sdk does not build against the lab's /opt/et, and there is no shared installUnchanged: still no /opt/gp-sdk on any host; it waits for one /opt/et (C18). | all three | blocks work | Open workaround | Nekko team RT2, RT6 |
| D11 | Kernel ELFs built on different hosts hash differently, but the code is identicalUnchanged (cosmetic): the toolchain in /opt/et is as on 25 Sep. | all three | cosmetic | Open workaround | Nekko team RT2 |
| D12 | A stray write from a kernel leaves no trace on the card or the hostUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep. | all four cards | corrupts results | Open workaround | Nekko team CF11 |
| D13 | Cycle counting traps: hpmcounter3 reads 128 short, the cycle CSR traps, and evict_va is asynchronousUnchanged: a card-side problem, and the cards, firmware and runtime are as on 25 Sep. | all four cards | corrupts results | Open workaround | Nekko team RT3, DI3 |
| D14 | Code on the lab hosts goes stale, or breaks, when it is updated in placeOne more instance, fixed and written down; the lab part is three lines in the onboarding page. | all three | wastes time | Open workaround | Nekko team DI1 |
| D15 | Long or remote jobs over Tailscale SSH die, kill themselves, or expose their command linesHit again on 27 Sep; the trap in our checklist was fixed that night (D22). Command lines are still readable by every user on all three hosts (30 Sep). | all three | wastes time | Open workaround | Nekko team AS2, DI1 |
| D16 | Our scripts/deploy-lab.sh used macOS tar flags and failed on LinuxHolds: scripts/deploy-lab.sh uses the portable tar wrapper. | our repository | blocks work | Fixed 25 Sep | nothing left |
| D17 | A coding agent told not to touch the cards ran a real measurement block on a shared cardSafeguards extended for the heat code; fail-closed dry runs still to do. | one lab host | corrupts results | Open workaround | Us U13; Nekko team PO2 |
| D18 | Our published notes and the lab-access page still carry the old diagnosesThe old diagnoses are corrected in the repository and the mirrored pages (25–27 Sep); the public accounts page's correction is prepared but not published; new errors moved to D23. | documentation | wastes time | Partly fixed | Us U7 |
| D19 | Two host-to-card copies on one stream move less than one, and the link gives no full-duplex gainNarrowed on 29 Sep by E55 (pre-registered, three cards): only two copies on one stream collide, moving 0.49 of one, while one copy on each of two streams moves 1.01, so the loss is per stream; a shared DMA read engine and the IOMMU are refuted, and the IOMMU-passthrough boot (U26) is no longer needed. Both directions at once give 1.07–1.08× (E50). Every card and root port run MaxPayload 256 B and MaxReadReq 128 B (28 Sep). | three cards | wastes time | New workaround | Nekko team DI5 |
| D20 | A small copy or an empty kernel costs hundreds of microseconds, mostly the runtime polling for completionE50: 556–566 µs for a waited empty kernel against 104 µs queued; 377–411 µs for a lone 4 KB copy. | all three | wastes time | New workaround | Nekko team RT5, RT2 |
| D21 | Documents the open drop cites but does not contain, so parts of the chip must be inferredThe chip diagram and the memory-level pages mark these parts “inferred”. | documentation | wastes time | New workaround | Nekko team DI6, CD1–CD6 |
| D22 | Our own checklist tells agents to run pgrep -af queue.sh over ssh, which always matches itselfFixed on 27 Sep at 23:34 and holding: AGENT.md §11 step 4 uses pgrep -af '[q]ueue.sh' and says why (commit a745199, pushed that night), and our card-behaviour notes and public newcomer brief carry the bracketed form. The 27 Sep update counted it as not committed yet. | our docs | wastes time | Fixed 27 Sep | nothing left |
| D23 | Our docs carry new wrong or sensitive statements: card 1 “firmware DVFS”, a root SSH route, a stale numpy notePartly corrected: card 1's rows in AGENT.md and our card-behaviour notes are right, but the lib.sh comment is not. Its correction was reverted on 28 Sep to keep lib.sh at the bytes our locked experiments use, so it still says aifoundry3 has no system numpy, and it waits for the owner's decision on that lock. add-lab-user.sh still names the root route; that is the owner's call (U14). | our repository | wastes time | New | Us U13, U14 |
| D24 | Our PCIe probe held a card's lock for 12.8–14.9 s per run, over the lab's 10 s ruleFixed on 28 Sep: run_pcie.sh releases the lock between sub-tests, and its first card run (aifoundry1's card 1, 23:56) took at most 1.84 s per sub-test, with no lock wait; the stale copies on aifoundry1 and aifoundry3 were replaced. | three cards | cosmetic | Fixed 28 Sep | nothing left |
| D25 | Reading the host CPU's energy needs root, so CPU-against-card energy comparisons rest on an assumed CPU powerThe host CPU's energy counter (/sys/class/powercap/intel-rapl:0/energy_uj) is readable by root only on all three hosts (checked 4 Oct), the kernel's default since a 2020 side-channel fix (CVE-2020-8694). So our 29 Sep comparison of CPU and card energy had to assume 125–251 W for the CPU. A policy choice, not a fault: the lab decides (MO8). | all three | wastes time | New | Nekko team MO8; then us (root) |