Ridge points: how much reuse the ET-SoC-1 needs at each level of memory
A roofline ridge point is the arithmetic intensity at which a kernel stops waiting for memory: peak compute divided by the bandwidth of the level the data comes from. Below it, that level sets the speed; above it, the arithmetic units do. For the ET-SoC-1's tensor unit at 600 MHz, fp32 work needs about 4 FLOP per byte loaded from the minion's own shire (its L2 cache or L2 scratchpad), 10 per byte from L3 or another shire's scratchpad, 130 per byte from DRAM, and about 790 per byte from host memory over PCIe (at the host-to-card rate timed on three cards on 27 September; 624 at the link's figure). fp16 needs twice as much and int8 eight times, because the tensor unit does 2× and 8× as many fp16 and int8 operations per cycle. One full-size TensorFMA already does 4 FLOP per byte it loads in fp32 (8 in fp16), so tiles held in the shire can just keep fp32 and fp16 busy, with no margin: with private tiles in each shire the matmul loop ran fp32 and fp16 at 529 cycles per op on all three lab cards, as fast as with shared tiles. Beyond the shire, and above all from DRAM, every byte has to be reused many times. From DRAM that takes C blocks of about 520 × 520 in fp32 and fp16, and about 1,040 × 1,040 in int8.
fp32 on the tensor unit at 600 MHz, with the bandwidth measured on the card (PCIe: the host-to-card rate timed on three cards on 27 September). Second line: fp16 and int8.
Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan on aifoundry2, aifoundry3 and aifoundry1 card 1. Of 27 claims tested here, this page counts 20 held and 7 corrected; the hub’s scoreboard, 15 “proven on the cards” and 12 “a test behind it failed”. The private-tile matmul settled what this page had left open (Own shire). The corrections are about a lone minion's loads: the smallest TensorLoad costs about 42 cycles, not 45; the DRAM rate of a lone minion, and so the minions it takes to saturate DRAM, depends on the kernel; "4 lines per round trip" was not confirmed; and the short probe's L2 rate varies between repeats. The check's energy test (RL-X4) found no verdict for L3 and another shire's scratchpad at the spec rates, or for TensorSend inside a shire on zeros (Not established), and the energy figures below are the energy manual's three-card values. Record: docs/reports/data/2026-09-25-claims-v3.
Terms used on this page
The compute cores are minions: small in-order RISC-V cores, each with two hardware threads (harts), an 8-lane vector unit and a tensor unit; 8 minions form a neighbourhood and 32 a shire, and 32 shires (1,024 minions) run kernels, joined by a mesh network-on-chip (NoC, 400 MHz on these cards) to 32 GB of LPDDR4X DRAM on 16 channels. These cards' firmware splits each shire's 4 MB of SRAM into a 512 KB L2 cache, a 1 MB slice of the chip-wide 32 MB L3 and a 2.5 MB L2 scratchpad that any shire can address, and each minion sets aside 3 KB of its 4 KB L1 data cache as an L1 scratchpad. A TensorLoad copies up to 1 KB into the L1 scratchpad; a TensorFMA multiplies an A tile from there by a B tile streamed through the TenB buffer (or also held in the L1 scratchpad) into a 16×16 C tile in the vector registers (int8 accumulates in a separate buffer, TenC); and a TensorSend moves data from one minion's registers to another's, inside a neighbourhood over the tree-shaped fast local network. More in the hub's glossary.
The compute ceilings
Each minion's vector unit has 8 lanes. Each lane has one FMA unit that does one fp32 or two fp16 multiply-adds per cycle, and two integer units that do four int8 multiply-adds each. A tensor instruction keeps those units busy for 512 cycles (fp32 and fp16) or 256 cycles (int8) per 16×16×K op, so a minion peaks at 16 fp32 FLOP, 32 fp16 FLOP or 128 int8 OP per cycle. The int8 rate is 8× fp32, not 4×, because the integer units are separate and there are twice as many. Only hart 0 of a minion may issue TensorFMA (hart 1 can only prefetch into the L2 scratchpad), and the vector unit's fmadd.ps runs on the same FMA units. So hart 1 and the vector unit add nothing to these peaks.
| Compute path | Per minion per cycle | 1,024 minions at 600 MHz | At 1 GHz | Best measured |
|---|
How the peaks compare with measurements
Measured: the matmul report's tensor loop with tiles in L2, 529 cycles per fp32 or fp16 op and 280 per int8 op, against 512 and 256 ideal, the same in every launch on all three cards. Vector: the best published vector-unit matmul, 2.94 TFLOP/s at 650 MHz (FOSDEM 2026). The vendor's 836 TOPS int8 for a six-chip card (139 per chip) is the same 128 OP per cycle on all 1,088 minions at 1 GHz. aifoundry3 is held at 600 MHz (a boot service sets its TDP to 0 W at every boot); aifoundry2's governor runs it at up to 800 MHz, the top of the firmware's frequency table, when the die is cool. Every lab measurement on this page is at 600 MHz; only the section on which ridge points move with the clock uses aifoundry2's faster launches.
Ridge points by level
Bandwidths are given per minion per minion-clock cycle, as a kernel running on all 1,024 minions sees them, and for the whole chip at 600 MHz. Measured is the median streaming rate of the earlier reports' probes with every minion reading its own data. Bytes per cycle come from the cycle counters, and GB/s from the 600 MHz launches only. The memory-hierarchy report uses the same launches. Spec is what the documented port and bank widths allow. Ridge = peak per minion-cycle ÷ bytes per minion-cycle.
Pick where the data comes from, a precision and the operations the kernel does per byte it fetches; the readout gives the fastest it can run. Speed is capped by the lower of the compute ceiling (horizontal) and the bandwidth line of the level (diagonal); the dots are the ridge points, labelled for the selected precision. The clock buttons apply the model in which ridge points move with the clock, measured on aifoundry2 only: the ceilings and the bandwidths inside a shire scale with the minion clock, DRAM and PCIe stay fixed, and L3 and another shire's scratchpad span a wedge between rising as measured and staying fixed. TensorSend between shires is drawn fixed: it is NoC-bound, not measured against clock, and spans six shire-to-shire patterns. The rings are kernels measured at 600 MHz: the matmul benchmark with tiles in L2 (fp32, fp16, int8), the same fp32 loop with every tile from DRAM, and the sparsity report's batch-1 fp32 layer with its weights in the scratchpads; the two below the ceilings sit on their bandwidth lines. The matmul rings are aifoundry2's first run and the layer aifoundry3's; the three-card check repeated each on all three cards (the tiles-in-L2 rings to 0.05%, the DRAM ring to 1%, the layer to 0.01 µs). The int8 matmul sits above the measured own-shire line (39 TOP/s at 16 OP per byte; it ran at 72) because its minions shared one tile pool (Own shire); with private tiles it ran on that line. It stays under the spec line. The tables below hold every number without the chart. A100 roofs overlays, dash-dotted, the A100's dense peaks (fp32 on the CUDA cores, fp16 and int8 on the tensor cores) and its three sustained bandwidths from the A100 table, with its ridge points on the selected precision's ceiling; the A100 lines do not move with the clock buttons.
Measured bandwidth and ridge points, as a table
| Level | fp32 | fp16 | int8 | Bandwidth |
|---|
Ridge points in FLOP (int8: OP) per byte at 600 MHz. Bandwidth: bytes per minion per cycle · the whole chip. For TensorSend rows, "bytes" are bytes one minion sends another.
- Inside a minion, nothing binds. The tensor unit reads its A operand from the L1 scratchpad at one request of at most 32 B per cycle. With B streamed through TenB, a full-size op keeps that port 50–75% busy. (In int8 with B also in the scratchpad it reaches 88%, and the op slows from 280 to 318 cycles, on all three cards.) fp32 and fp16 keep C in the vector registers, whose three 32 B read ports per cycle are sized for one three-operand instruction per cycle. int8 accumulates in the separate 1 KB TenC.
- Own shire. The probes stream 1 KB TensorLoads, two in flight per minion. With all 1,024 minions reading their own data, they level off at 4.0 B per cycle each (128 B per cycle per shire), from the L2 cache and the scratchpad alike, on
aifoundry2. The 23 September reruns of the same probe at a pinned 600 MHz gave 4.00 B per cycle for both onaifoundry3as well; only the sparsity report's shorter L2 probe (2,000 loads per minion) read less, 3.3 (slowest minion). The three-card check found that short probe varying from 2.7 to 3.9 B per cycle between repeats, while the same probe with 20,000 and 200,000 loads read 3.99 and 4.00 on all three cards. The shire's four cache banks can return 256 B per cycle, twice the 4.0 rate. The int8 matmul loop consumed 7.3 B per cycle per minion, but all 1,024 minions read the same 32 KB of tiles, so some lines probably came from the L2's read buffers rather than its banks. With private tiles in each shire the same loop drew 4.0 B per cycle (512 cycles per int8 op) on all three cards, so the shared pool accounts for the 7.3. So 4 FLOP/B is what a simple streaming kernel gets, and no kernel reading private data, the private-tile matmul included, has yet approached the banks' spec ridge of 2. - Other shires. L3 lines are spread over all 32 shires, so almost every L3 hit crosses the mesh, just like a read from another shire's scratchpad. Both measured 0.96–0.98 TB/s, or 1.6 B per minion-cycle. The remote rate is for one pattern: each shire reading the shire 16 IDs away, 2.1 hops on average. If many shires read one shire's scratchpad, that shire's four mesh ports would allow only about 0.1 TB/s. At the NoC's 400 MHz, the ports of all 32 shires together would allow about 3.3 TB/s if each moved one 64 B beat per NoC cycle. No document states that rate for the shire ports; the datasheet's 32 GB/s per memory-shire NoC port is one 64 B beat per cycle at the 500 MHz design NoC clock. The mesh links, whose width is undocumented, could limit traffic spread over all shires first.
- DRAM. Measured at 76 GB/s on both cards (75.8–76.0 GB/s in the 23 September reruns at 600 MHz; the sparsity report's shorter probe read 72 GB/s on
aifoundry3, and 71–73 GB/s on all three cards in the three-card check, where the same probe with ten times as many loads read 75.6–75.9 GB/s). For comparison, 16 channels give 119 GB/s at the 3,733 MT/s implied by the firmware's 933 MHz DDR clock, and 136.5 GB/s at the datasheet maximum of 4,266 MT/s. The datasheet prints 133 Gbytes/s, which is 136,512 MB/s divided by 1,024. The 128,000 MB/s the runtime reports is a constant in the service processor's firmware, not a measurement. A third party's estimate of about 88 GB/s on another card was calculated from single-memory-shire rates, not measured chip-wide. - Host memory. PCIe Gen4 x8 carries 15.75 GB/s per direction before protocol overhead. Timed on 27 September (the PCIe link page), a 256 MB copy's DMA moves 12.46–12.60 GB/s on the three cards from host to card and 10.41–10.54 GB/s back. A program's copy, staged through the runtime's bounce buffer, gets only 5.2–7.8 GB/s from host to card, set by each host's own memcpy. The ridge above uses the host-to-card DMA rate pooled over the cards, 12.5 GB/s.
- TensorSend moves data between minions' registers rather than fetching it, and each rate belongs to the pattern the on-chip communication report measured: pairs, rings, or shire to shire. It is not the capacity of a link. Across the mesh, a shire manages only 3–5 GB/s, limited by how many packets a sender keeps in flight.
Spec bandwidths and ridge points, as a table
| Level | Bandwidth, spec | Spec ridge, fp32 · fp16 · int8 |
|---|
The same levels at what the documented port and bank widths allow, in the same units.
The same numbers for an A100
| Level | Bandwidth, sustained | fp32 | fp16 tensor | int8 tensor |
|---|
How the A100 numbers were computed
Ridge points against sustained bandwidth, each ET-SoC-1 level next to its closest A100 counterpart. The A100's peaks are 19.5 TFLOP/s fp32 on the CUDA cores, 312 TFLOP/s fp16 and 624 TOP/s int8 on the tensor cores (dense). Its bandwidths are the sustained values in docs/reports/sources/2026-09-18-a100-memory-hierarchy.md (HBM2 1.40 TB/s). The ET-SoC-1's fp32 runs on its tensor unit.
The A100's tensor cores outrun its bandwidth by more, so on chip its fp16 and int8 ridge points are higher than the ET-SoC-1's. From DRAM the fp16 ridges are close (223 against 259). In int8 the ET-SoC-1 needs 2.3× more, and in IEEE fp32, which the A100 runs on its CUDA cores, about 9× more (130 against 14). With fp32 inputs rounded to TF32 on its tensor cores (156 TFLOP/s dense), the A100's HBM ridge would be about 111.
Which ridge points move with the clock
Only aifoundry2's sessions show how bandwidth follows the clock, and this section rests on its two: aifoundry3 is held at 600 MHz (a boot service sets its TDP to 0 W at every boot), while aifoundry2's governor moved it between 600 and about 800 MHz from launch to launch. Levels inside the shire run on the minion clock, because the shire cache is clocked with the minions. The L2 and scratchpad bandwidths rose exactly with the clock, so their ridge points hold at any clock. L3 and remote scratchpad reads go through the NoC, which the firmware keeps at 400 MHz. Over launches up to about 740 MHz, their bandwidth rose only 0.65× and 0.33× as fast as the minion clock. DRAM bandwidth did not change at all, so its ridge point grows in proportion to the minion clock. PCIe should behave the same way, but its rate has been timed only at 600 MHz.
| fp32 ridge (FLOP/B) | 600 MHz | 800 MHz | 1,000 MHz |
|---|
Method and caveats for the clock model
fp16 is 2× and int8 8× every entry. 600 MHz is measured on both cards; the 800 MHz (the top of the firmware's table) and 1,000 MHz (the shire design clock) columns apply the clock scaling seen on aifoundry2. Only the minion clock changes. The NoC stays at the lab cards' 400 MHz and DDR at 933 MHz; at the NoC's 500 MHz design clock the L3 and remote-scratchpad ridges would be lower. For those two, the lower end assumes their bandwidth keeps rising with the clock as it did between 600 and about 740 MHz, and the upper end assumes it stays fixed. TensorSend inside a neighbourhood or a shire runs on the minion clock, so its ridges should hold at any clock; that is inferred from the clock domain, not measured, since TensorSend has only been run at 600 MHz. Between shires it crosses the 400 MHz NoC.
Every launch, plotted against clock
One dot per launch of the memory-hierarchy report's streaming probes on aifoundry2 (18 September, both sessions), plotted at the minion clock its cycle counts imply. y = 1 is that level's own median GB/s at 600 MHz; a level whose bandwidth simply followed the clock would sit on the dashed line. L1, L2 and the own scratchpad (inside the shire, on the minion clock) track it; L3 and another shire's scratchpad (crossing the fixed-frequency NoC) rise more slowly; DRAM does not move at all.
What it takes to reach them
Matrix multiplication
How each reuse ridge is derived
- One op. A 16×16×K TensorFMA reads a 1 KB A tile from the L1 scratchpad and a 1 KB B tile streamed through TenB. It does 8,192 FLOP in fp32, 16,384 in fp16 or 32,768 OP in int8. With both tiles loaded fresh each time, that is 4, 8 or 16 per byte: exactly the own-shire ridge for fp32 and fp16, and half of it for int8. The matmul benchmark's 97% and 91% of peak came from one shared tile pool (Own shire); with private tiles the same loop ran at 529 cycles per op in fp32 and fp16 (97%) and 512 in int8 (50%) on all three cards: a kernel with its own tiles in the shire sits exactly at the fp32 and fp16 ridge, with no margin, and int8, at half its ridge, runs at the own-shire streaming rate.
- The shire. A shire's 32 register-resident C tiles form a 64 × 128 block. A block of m × n outputs held while K streams past does 2mnK operations for every e·K·(m + n) bytes it reads (e = bytes per element), an intensity of H/e with H = 2mn/(m + n), the harmonic mean of m and n. Here H = 85: 21, 43 and 85 operations per byte. That is well past the banks' spec ridge of 2, 4 and 16 if a cooperative TensorLoad reads each line once for every minion that needs it (A shared by the 8 minions of a neighbourhood, B across the 4 neighbourhoods: 384 B of bank reads per minion-op instead of 2 KB). It also beats the measured 10, 20 and 80 from L3 or another shire, with little to spare in int8, if each byte enters the shire once. Cooperative loads have not been measured on these cards, and they cannot read L3-homed addresses, so L3 data must first be captured by the shire's L2 or copied into its scratchpad.
- DRAM. These figures assume each DRAM byte is read once for the whole chip. The block needs H ≥ e × ridge: about 520 in fp32 and fp16, and about 1,040 in int8. All 1,024 register tiles together form a 512 × 512 block (1 MB of C), which gives 128, 256 and 512 per byte. That is 99% of what fp32 and fp16 need and half of what int8 needs. Larger blocks have to keep C in the scratchpads and swap it through the registers. The 80 MB of scratchpad (the default partition) would hold C for a block of about 4,580 × 4,580 in fp32, but nobody has measured what the swapping costs. int8 is harder: its accumulator, TenC, cannot be loaded from memory, so partial sums would have to be added with vector instructions.
How to read this chart
Fraction of the tensor unit's peak at 600 MHz that the measured bandwidth of each level allows (the roofline above, read along one diagonal). Blocked matmul: x is H, the harmonic mean of the C block's sides (a square block's side); the intensity is H/e per byte. L3 and another shire's scratchpad are read by one shire, whose 32 register tiles hold at most a 64 × 128 block (H = 85), so their lines stop there; DRAM assumes each byte is read once for the whole chip, with C in all 1,024 minions' registers up to H = 512 and swapped through the scratchpads (unmeasured) up to about 4,580. Streaming weights: x is the batch N; the weights give 2N/e operations per byte, and inside the own shire each op also loads a fresh activation tile, so that line flattens at N = 16. The ticks are the smallest compute-bound blocks and batches, as in the table below.
Inference that streams its weights
A layer that reads its weights once per batch of N inputs does 2N operations per weight, or 2N/e per byte (e = bytes per element). It is compute-bound when that reaches the ridge, so it needs N ≥ e × ridge / 2:
| Weights come from | fp32 | fp16 | int8 |
|---|
Caveats on the batch table
These are the smallest batches that are compute-bound at the measured bandwidth. The own-shire row also counts the activations. A minion holds one 16×16 C tile, so each op loads a fresh activation tile as well as a weight tile. Its intensity therefore tops out at 4, 8 and 16 per byte at batch 16, and int8 never reaches its 32 OP/B own-shire ridge. The other rows assume the activations stay in the shire and each weight byte reaches the shire once. On A0 silicon, fp32 and fp16 ops of 1–4 rows with B from TenB must be padded to 5 rows (errata 1.29 type D).
A worked example: the batch-1 GEMV
At batch 1 a layer does 0.5 FLOP per byte in fp32, so it is memory-bound at every level. The sparsity report's batch-1 layer, with its weights in the scratchpads and its partial sums added by the host, ran at 1.24 TFLOP/s (6.8 µs per 1,024 × 4,096 layer: the row "Every row loaded, compute only" in that report's layer table). That is the scratchpad bandwidth times 0.5, within 1%. With the on-chip reduction the layer takes 7.3 µs, 1.15 TFLOP/s. Both times repeat to within 0.01 µs on all three cards.
Other limits that act like ridge points
- A lone minion. A minion keeps at most 4 line requests in flight, shared by its two TensorLoad paths. How many a lone minion actually has in flight was not settled: from DRAM, where a 1-line load (73 cycles) is about one round trip, its 16-line loads took 749–751 cycles on all three cards in two of the three-card check's three passes, which implies about 3.1 requests in flight with one of the two paths in use (32 × 73 ÷ 750), and 1,132 cycles in the first pass on every card (2.1). Streaming its own data alone, one minion measured 6.4 B per cycle from L2 or the scratchpad (6.38 and 6.40 on all three cards) and 1.4 from DRAM (1.37 in two of three passes on all three cards, 0.90 in the first). This is not a hard ceiling: the int8 matmul loop drew 7.3 B per cycle per minion, all minions reading one shared tile pool. The lone-minion DRAM rate depends on the kernel: the energy catalogue's TensorLoad loop (
enercat, 1 KB loads over 64 MB) gets 0.91 B per cycle alone and 0.88 per minion with 32 minions streaming together, on all three cards. So it takes at least about 93 minions streaming 1 KB loads at once to reach the chip's 76 GB/s from DRAM at the sparsity probe's rate, at least about 140 at that loop's, and more once queueing raises the latency.
Three more limits: small loads, launch overhead, the vector unit
- Small loads. Every TensorLoad costs at least about 42 cycles, even with every line masked off (42–43 cycles on all three cards in the three-card check, from L2 or the scratchpad; 45–47 in the first run), and about 65 for up to 4 lines when all minions load at once. So loads of fewer than 4 lines cannot reach the 4 B per cycle streaming rate: a 1-line load gets about 1.0 B per cycle. 4-line loads still reach 3.9. These last three repeat on all three cards.
- Kernel launch. In the matmul runs, a launch from the host took 0.43 ms (median; 0.35–0.69 ms) on
aifoundry2; in the three-card check the medians were 0.42–0.48 ms onaifoundry2, 0.32–0.38 ms onaifoundry3and 0.43–0.47 ms on aifoundry1 card 1 (0.245–0.80 ms over all launches). That launch time is as long as 4.2 GFLOP takes at the fp32 peak, so a kernel with less work than that spends more time launching than computing. - The vector unit's limit is not pinned down. A
flw.psloads 32 B and anfmadd.psdoes 16 FLOP, and a minion issues at most one instruction per cycle for its two harts together. So with I FLOP per loaded byte, at most a fraction 2I/(2I + 1) of the issue slots can be FMAs. Measured code stays well below that bound. With both harts doing nothing but loading, the energy manual's catalogue loop (64 loads per loop iteration) issued aflw.psevery 1.4 minion-cycles, almost as fast as it issued an integeradd, so the L1 delivered at least 23 B per minion-cycle (14.2 TB/s) and the load rate binds first only below about 0.69 FLOP per byte. The memory-hierarchy probe's loop (8 per iteration) issued one every 3.2 cycles, 10 B per minion-cycle: a limit of that loop, not of the L1. The best published vector matmul has 63% FMAs in its instruction mix, yet reaches only 4.4 FLOP per minion-cycle, 28% of peak: it issues about 0.44 instructions per cycle. The spec ridges against the L1 cache (0.5 FLOP/B, 32 B per cycle) and the registers (0.17, 96 B per cycle) are never reached.
Energy balance points
Time is one budget and energy is another. Divide each level's energy per byte by the tensor unit's energy per FLOP, both measured as card power above idle. The result is the intensity at which moving the data costs as much energy as the arithmetic. On random-normal operands held in the L1 scratchpad, the fp32 tensor unit costs 2.89 pJ per FLOP above idle [2.62–3.01] and int8 0.147 pJ per OP (energy manual §3.2, at 600 MHz; four runs on each of the three cards). By card, fp32 reads 2.96, 2.68 and 3.03 pJ per FLOP (aifoundry2, aifoundry3, aifoundry1 card 1). These are the ablation's registered values, which carry a launch-temperature offset whose size depends on how the die sensor's whole-degree reading at launch maps to the die temperature (amendment C2): aifoundry2's read from 0.6 W high to 0.2 W low, aifoundry3's 0.9–1.4 W low and aifoundry1 card 1's from 0.7 W high to 0.05 W low. So at the same die temperature aifoundry3's fp32 energy is 0.96–0.97 of aifoundry2's, not the registered 0.91 (the energy manual's catalogue, measured separately, gives 0.972), and the pooled value moves by at most 2%. The balance points below use the registered values. The data matters as much as the level: on all-ones operands fp32 costs 1.11 pJ per FLOP and on zeros 0.21, so every balance point below moves by up to 14× with the operands.
Removing the cost of keeping the cores awake
Both sides include the cost of keeping the cores awake (about 1.5–2.1 W for the whole chip; energy manual §2), which a kernel doing both pays once. Taking it out of both lowers the random-data balance points from the shire's own scratchpad to DRAM by about 9–29% (TensorSend's by more, since this cost is most of its power over idle) and changes none of the conclusions below; on zeros it is most of the tensor unit's 0.21 pJ per FLOP. On zeros the launch-temperature offset is as large as the signal: the tensor unit on zeros draws 1.85, 0.75 and 1.15 W over idle as registered, so the cards' zeros figures (0.20, 0.08 and 0.13 pJ per FLOP) cannot be compared with each other.
How to read this chart
One row per level, fp32 on the tensor unit: the circle is the energy balance point (byte energy ÷ FLOP energy, pooled over the three cards), the whisker every card's 99% range, the diamond the time ridge. A verdict is drawn only where every card's range clears the diamond; left of it, arithmetic costs more energy than moving the data. On zeros read the rows as a direction, not a value; the L1 row is in the table.
The energy roofline chart
All the energy figures, as a table
| Where the bytes come from | pJ/B | fp32 balance | int8 balance | fp32 time ridge |
|---|
Balance points and the time ridge are in FLOP (int8: OP) per byte. Byte energies are the energy manual's §4.2 (600 MHz, three cards, six passes each; the scratchpad rows on random data, on zeros 2.25 and 5.10 pJ/B) and §5 (TensorSend, 1 KB messages; the last row six shire-to-shire patterns). Only the own scratchpad differs between the cards beyond noise (aifoundry1 card 1 reads higher), so its row gives each card's value, aifoundry2 · aifoundry3 · aifoundry1 card 1, each balance point against its own card's FLOP energy. The int8 column pools the three cards' int8 energy (0.157, 0.131 and 0.154 pJ per OP, registered values). The L1 row sets a 32 B vector load against the vector unit's fmadd.ps (3.49 pJ per FLOP, the manual's §4.1 catalogue on random data), at that loop's 14.2 TB/s (the memory-hierarchy probe's slower loop: §4.2).
On random-normal operands, every fp32 energy balance point lies below the matching time ridge point at the measured bandwidths. So a compute-bound kernel on random-like data spends more energy on arithmetic than on moving data. The operands can change this. On all-ones operands the own scratchpad and TensorSend still lie below their time ridges, but the L2 cache, another shire's scratchpad, L3 (9.4 against 10) and DRAM (89 against 130, with the byte energy also measured on all-ones data) are not clear of theirs on every card (not established). On zeros the tensor unit adds little more power than cores that are merely awake, and the shire's own scratchpad, the L2 cache, L3 and DRAM lie well above their time ridges, even with byte energies also measured on zeros (§4.1: DRAM 434 against 130, own scratchpad 9.6 against 4.0). For zero-heavy operands, moving the bytes costs more energy than the arithmetic even when the kernel is compute-bound. Treat the zeros figures as a direction, not a value.
Checking against random-data byte energies
The §4.2 probes of the L2 cache, L3 and DRAM read whatever the buffers held, not random data. On random data a TensorLoad costs 4.21 pJ/B from the shire's own scratchpad and 129 pJ/B from DRAM (§4.1), which puts those balance points at 1.5 and 45 FLOP/B. That does not change the conclusion at the measured rates. At the spec rates 1.5 is below the shire banks' spec ridge of 2 on every card, and 45 is below DRAM's spec ridge of 72–82, but on aifoundry1 card 1 the DRAM figure's 99% range reaches 84, so that comparison is not established on every card.
Not established: comparisons too close to call
Not established
In these comparisons at least one card's 99% range reaches the ridge (the chart's tooltips give each card's range), so the page draws no verdict for them.
- Spec rates, random-normal operands. The three-card check tested L3 and another shire's scratchpad against the spec ridge of 3.0, itself an assumption (one 64 B beat per NoC cycle), under its registered rule (RL-X4: byte energies over every pass of 23 and 26 September, the FLOP energy over the ablation's four runs per card, a verdict only when every card's 99% range lies on one side). L3's balance point read 3.97 [2.39–5.61] FLOP/B on
aifoundry2, 4.63 [2.72–6.59] onaifoundry3and 6.08 [3.48–8.84] on aifoundry1 card 1, and another shire's scratchpad 2.52 [1.38–3.69], 2.79 [1.63–3.97] and 3.19 [0.79–5.74]. Only aifoundry1 card 1's L3 range clears the ridge (above it); the others overlap it, so neither comparison has a verdict. On this page's pooled figures L3 sits at 3.6 and another shire's scratchpad (random fill) at 2.3, and the L2 cache (1.1 against 2.0) and DRAM (46 against 82) are not clear of theirs on aifoundry1 card 1. - All-ones operands. The L2 cache (2.2 against 4.0), another shire's scratchpad (random fill; 6.0 against 10), L3 (9.4 against 10) and DRAM (89 against 130) are not clear of their time ridges on every card. Against the spec ridges, the own scratchpad (2.4 against 2.0), the L2 cache (2.2 against 2.0), another shire's scratchpad (6.0 against 3.0) and DRAM (89 against 82) overlap theirs; only L3 (9.4 against 3.0) lies above its.
- Zeros. TensorSend inside a shire (10.0 against 8.8) and between pairs (3.2 against 3.3) overlap their time ridges, as do another shire's scratchpad (filled with zeros; 32 against 10, not clear of it on aifoundry1 card 1) and TensorSend between shires over its six patterns (57–85 against 63–112). The three-card check's registered test of TensorSend inside a shire (RL-X4) read 10.4 [7.4–15.9] FLOP/B on
aifoundry2, 25.5 [18.7–37.6] onaifoundry3and 19.7 [14.4–28.5] on aifoundry1 card 1 against 8.8: above it onaifoundry3and aifoundry1 card 1, not clear of it onaifoundry2, so no verdict; and on zeros the launch-temperature offset above is comparable to each card's FLOP side itself.
Caveats
Caveats in full
- The measured bandwidths are what the probes reached with 1 KB TensorLoads, two in flight. The spec values are what the documented port and bank widths allow. Neither is a proven ceiling. The spec figure for the mesh assumes one 64 B beat per NoC cycle per port.
- These are read ridge points. Writes were measured later (energy manual §4.1, three cards): tensor stores reach 1.23 TB/s into the shire's own scratchpad, half the read rate, and 76 GB/s to DRAM, the same as reads; stores through the L1 reach DRAM at only 27 GB/s. A kernel that writes as much as it reads needs more reuse than these tables show.
- The levels beyond the shire use the 600 MHz launches of the 18 September runs on
aifoundry2, whose governor moved the clock; the 23 September reruns at 600 MHz (pinned onaifoundry3, held there onaifoundry2by a warm die) reproduce every level on both cards to within 0.3%.aifoundry3also supplied the sparsity report's TensorLoad-size and batch-1 measurements, and the three-card check (26 September) repeated them, the matmul and the private-tile matmul onaifoundry2,aifoundry3and aifoundry1 card 1 (firmware 1.2.0). - The port widths come from the documentation and RTL of core-et's Erbium branch (commit
b38a1a3): the same Minion core lineage in a later MCU-class configuration, not the taped-out ET-SoC-1. A0 silicon may differ in details.
Reproduce
Reproduce this
python3 scripts/ridge-points.py --embed docs/reports/2026-09-18-et-soc1-ridge-points.html
python3 scripts/paste-chartkit.py docs/reports/2026-09-18-et-soc1-ridge-points.html
The script reads the raw data of the earlier reports (docs/reports/data/), the 23 September reruns of the memory probes, the on-chip communication report's embedded JSON and the energy manual's data (docs/reports/data/2026-09-23-energy-manual/manual.json). It prints the numbers behind the tables, the charts and the figures in the text, and writes the tables' and charts' data into this page. The spec constants in the script each name their source. Its inputs rebuild from committed data with the commands in docs/findings/04-artifacts.md (A19); only the raw probe and matmul runs needed a card. The second command pastes in the shared chart toolkit the charts are drawn with.
Versions. First published 24 September with energy balance points from the 18 September memory-hierarchy energies (clock free) and 1.88 pJ per FLOP on small-integer operands. Recomputed the same day from the energy manual, which moved the DRAM balance point from 79 to 42 FLOP/B and the L2 cache's below its spec ridge. Revised 25 September: the L1 row now sets a vector load against the vector unit's fmadd.ps (balance 0.15, was 0.27 against the tensor unit); the energy conclusion is stated for random-normal operands, with the reversal on zeros; the rerun check reads the passes the energy manual pools (0.3%, was 0.2%); the vector unit's L1 load rate is the energy manual catalogue loop's 23 B per minion-cycle (below about 0.69 FLOP per byte the load rate binds, was 1.6 from the memory-hierarchy probe's slower loop at 10 B); and the page gained interactive roofline, energy and reuse charts. 25 September (version 3): checked claim by claim against both cards (record: docs/reports/data/2026-09-25-claims-v3); a figure with no card named holds on both. The own scratchpad's and L2 cache's energy per byte are now per card, the int8 balance points use aifoundry2's own byte energies (DRAM 777, was 771), and an energy verdict needs both cards' 99% ranges to clear the ridge, which moves L3 at the spec rates (was "lies above") and nine other comparisons the chart drew to Not established; the all-ones DRAM comparison uses bytes measured on all-ones data (89, was 109). 26 September: the three-card check (aifoundry2, aifoundry3, aifoundry1 card 1) added to the text: the private-tile matmul (it settles the own-shire question), a lone minion's loads (floor about 42 cycles, was 45; its DRAM rate per kernel; lines in flight not settled), the short probes' ranges and the launch time per card; and the energy balance points recomputed from the energy manual's three-card data (fp32 and int8 on each card, the launch-temperature offset stated; the levels re-run on three cards, the scratchpad rows filled with random data, only the own scratchpad differing by card), which takes the L2 cache and DRAM at the spec rates, and the L2 cache and DRAM on all-ones operands, out of the verdicts (Not established, with the check's RL-X4 values). 27 September: bandwidth against the clock and the energy roofline charted (the roofline per card); the scratchpad-held block is about 4,580 on a side (was 4,500). Later that day the host link's ridge moved from the link's figure (15.75 GB/s, 624 FLOP/B in fp32) to the host-to-card DMA rate timed on three cards (the PCIe link page; about 12.5 GB/s, 786 FLOP/B), drawn solid on the roofline with the link's figure among the spec lines. 28 September: the review's cuts (the shared-tile explanation once, the energy chart's caption and table note); the note's counts given by both rules.
Related reports
- Matmul efficiency — the tensor loop whose measured rates the compute peaks here are checked against.
- Memory hierarchy — the 18 September bandwidth probes behind the ridge table, with latencies and working-set sizes.
- On-chip communication — the TensorSend rates, and where the shires sit on the mesh.
- Sparse compute — the TensorLoad-size measurements behind "Other limits", and the batch-1 layer.
- The energy manual — §3.2 energy per multiply-add, §4 energy per byte read or written, §5 energy per byte sent: the inputs of the energy balance points.
- The Horace experiment — why the tensor unit's energy per FLOP depends on the data it multiplies.
- The PCIe link and the launch path — the host link timed on three cards: the host-to-card rate behind the PCIe ridge, and what a program's staged copy gets.
- Hand it to the next shire — a chain of stages that hands each stage's output to another shire's scratchpad (the next in ID order, 3.5 mesh hops away on average) instead of DRAM, 12.3× faster: the remote-scratchpad ridge in practice.
Sources
Sources and citations
- Minion VPU Specification (core-et docs, Erbium branch): §2 (8 lanes, one FMA unit and two int8 units per lane), §2.2.2 (register-file ports), §2.2.3.8 Table 6 (four int8 products per integer unit), §2.3.6.1 (tensor micro-op sequencing). ET Programmer's Reference Manual (aifoundry-org/et-man): §8.3.1 (L1 scratchpad mode), §9.1–9.4 (tensor op shapes and operands).
- CORE-ET Shire Cache Specification: §1 (4 banks, one 64 B line per cycle each; 512-bit NoC ports), §1.4.3 (L3 lines interleaved over all shires), §2.3–2.4 (four ports to the L3 mesh). CORE-ET Neighborhood MAS: §4.3–4.7 (neighbourhood buses, fast local network), Table 6 (shire and minion clocks). Minion DCache Description: §2.3.1 (256-bit scratchpad read port to the vector unit), §3.2 (L1 ports), §3.9.6 (TensorLoad requests in flight). CORE-ET Minion Shire Description: Table 2 (1 GHz shire and 500 MHz NoC design clocks). All in core-et.
- ET Preliminary Datasheet Rev 1.0 (in aifoundry-org/et-man): §1 and §7 (LPDDR4X, PCIe Gen4 x8), §2.1.1 (single issue, two harts), §7.1.1 (32 GB/s per memory-shire NoC port). Ditzel, Hot Chips 33 (2021), and Ditzel et al., IEEE Micro 42(3), 2022 (128 int8 OP per cycle per minion; 836 TOPS for six chips at 1 GHz, i.e. 139 per chip; HC33: 137 GB/s).
- et-platform service-processor firmware (
device-bootloaders/src/ServiceProcessorBL2):common/main.c(NoC PLL at 400 MHz),driver/mem_controller.c(933 MHz DDR),include/mem_controller.h(the reported 128,000 MB/s),services/thermal_pwr_mgmt.c(minion frequency table, 300–800 MHz). - P. Cawley, "Zero to matmul with the ET-SoC-1", FOSDEM 2026 (vector-unit matmul, 2.94 TFLOP/s at 650 MHz). M. Chang (marty1885), "Investigating the ET-SoC-1 NoC", 2026, also on the AI Foundry blog (about 88 GB/s from DRAM, calculated from the best rate against a single memory shire).
- A100: NVIDIA A100 datasheet (dense peaks, TF32 156 TFLOP/s);
docs/reports/sources/2026-09-18-a100-memory-hierarchy.md(sustained shared-memory, L2 and HBM bandwidth).