Ridge points: how much reuse the ET-SoC-1 needs at each level of memory

Ridge points computed 18 September 2026; energy balance points added and published 24 September 2026 · aifoundry2, aifoundry3 and aifoundry1 card 1 at a 600 MHz minion clock (NoC 400 MHz, DDR 933 MHz); a figure with no card named holds on each card measured · arithmetic on the 18 September measurements, the energy manual (its three-card data of 26 September) and the PCIe link page (27 September), no new runs · code: scripts/ridge-points.py · part of the ET-SoC-1 measurement reports

A roofline ridge point is the arithmetic intensity at which a kernel stops waiting for memory: peak compute divided by the bandwidth of the level the data comes from. Below it, that level sets the speed; above it, the arithmetic units do. For the ET-SoC-1's tensor unit at 600 MHz, fp32 work needs about 4 FLOP per byte loaded from the minion's own shire (its L2 cache or L2 scratchpad), 10 per byte from L3 or another shire's scratchpad, 130 per byte from DRAM, and about 790 per byte from host memory over PCIe (at the host-to-card rate timed on three cards on 27 September; 624 at the link's figure). fp16 needs twice as much and int8 eight times, because the tensor unit does 2× and 8× as many fp16 and int8 operations per cycle. One full-size TensorFMA already does 4 FLOP per byte it loads in fp32 (8 in fp16), so tiles held in the shire can just keep fp32 and fp16 busy, with no margin: with private tiles in each shire the matmul loop ran fp32 and fp16 at 529 cycles per op on all three lab cards, as fast as with shared tiles. Beyond the shire, and above all from DRAM, every byte has to be reused many times. From DRAM that takes C blocks of about 520 × 520 in fp32 and fp16, and about 1,040 × 1,040 in int8.

Own shire: L2 or scratchpad
4 FLOP/B
Other shires: L3 or a remote scratchpad
10 FLOP/B
DRAM
130 FLOP/B
Host memory over PCIe
786 FLOP/B

fp32 on the tensor unit at 600 MHz, with the bandwidth measured on the card (PCIe: the host-to-card rate timed on three cards on 27 September). Second line: fp16 and int8.

Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan on aifoundry2, aifoundry3 and aifoundry1 card 1. Of 27 claims tested here, this page counts 20 held and 7 corrected; the hub’s scoreboard, 15 “proven on the cards” and 12 “a test behind it failed”. The private-tile matmul settled what this page had left open (Own shire). The corrections are about a lone minion's loads: the smallest TensorLoad costs about 42 cycles, not 45; the DRAM rate of a lone minion, and so the minions it takes to saturate DRAM, depends on the kernel; "4 lines per round trip" was not confirmed; and the short probe's L2 rate varies between repeats. The check's energy test (RL-X4) found no verdict for L3 and another shire's scratchpad at the spec rates, or for TensorSend inside a shire on zeros (Not established), and the energy figures below are the energy manual's three-card values. Record: docs/reports/data/2026-09-25-claims-v3.

Terms used on this page

The compute cores are minions: small in-order RISC-V cores, each with two hardware threads (harts), an 8-lane vector unit and a tensor unit; 8 minions form a neighbourhood and 32 a shire, and 32 shires (1,024 minions) run kernels, joined by a mesh network-on-chip (NoC, 400 MHz on these cards) to 32 GB of LPDDR4X DRAM on 16 channels. These cards' firmware splits each shire's 4 MB of SRAM into a 512 KB L2 cache, a 1 MB slice of the chip-wide 32 MB L3 and a 2.5 MB L2 scratchpad that any shire can address, and each minion sets aside 3 KB of its 4 KB L1 data cache as an L1 scratchpad. A TensorLoad copies up to 1 KB into the L1 scratchpad; a TensorFMA multiplies an A tile from there by a B tile streamed through the TenB buffer (or also held in the L1 scratchpad) into a 16×16 C tile in the vector registers (int8 accumulates in a separate buffer, TenC); and a TensorSend moves data from one minion's registers to another's, inside a neighbourhood over the tree-shaped fast local network. More in the hub's glossary.

The compute ceilings

Each minion's vector unit has 8 lanes. Each lane has one FMA unit that does one fp32 or two fp16 multiply-adds per cycle, and two integer units that do four int8 multiply-adds each. A tensor instruction keeps those units busy for 512 cycles (fp32 and fp16) or 256 cycles (int8) per 16×16×K op, so a minion peaks at 16 fp32 FLOP, 32 fp16 FLOP or 128 int8 OP per cycle. The int8 rate is 8× fp32, not 4×, because the integer units are separate and there are twice as many. Only hart 0 of a minion may issue TensorFMA (hart 1 can only prefetch into the L2 scratchpad), and the vector unit's fmadd.ps runs on the same FMA units. So hart 1 and the vector unit add nothing to these peaks.

Compute pathPer minion per cycle1,024 minions at 600 MHzAt 1 GHzBest measured
How the peaks compare with measurements

Measured: the matmul report's tensor loop with tiles in L2, 529 cycles per fp32 or fp16 op and 280 per int8 op, against 512 and 256 ideal, the same in every launch on all three cards. Vector: the best published vector-unit matmul, 2.94 TFLOP/s at 650 MHz (FOSDEM 2026). The vendor's 836 TOPS int8 for a six-chip card (139 per chip) is the same 128 OP per cycle on all 1,088 minions at 1 GHz. aifoundry3 is held at 600 MHz (a boot service sets its TDP to 0 W at every boot); aifoundry2's governor runs it at up to 800 MHz, the top of the firmware's frequency table, when the die is cool. Every lab measurement on this page is at 600 MHz; only the section on which ridge points move with the clock uses aifoundry2's faster launches.

Ridge points by level

Bandwidths are given per minion per minion-clock cycle, as a kernel running on all 1,024 minions sees them, and for the whole chip at 600 MHz. Measured is the median streaming rate of the earlier reports' probes with every minion reading its own data. Bytes per cycle come from the cycle counters, and GB/s from the 600 MHz launches only. The memory-hierarchy report uses the same launches. Spec is what the documented port and bank widths allow. Ridge = peak per minion-cycle ÷ bytes per minion-cycle.

Roofline explorer: where does the data come from, and how fast can the kernel run?

Pick where the data comes from, a precision and the operations the kernel does per byte it fetches; the readout gives the fastest it can run. Speed is capped by the lower of the compute ceiling (horizontal) and the bandwidth line of the level (diagonal); the dots are the ridge points, labelled for the selected precision. The clock buttons apply the model in which ridge points move with the clock, measured on aifoundry2 only: the ceilings and the bandwidths inside a shire scale with the minion clock, DRAM and PCIe stay fixed, and L3 and another shire's scratchpad span a wedge between rising as measured and staying fixed. TensorSend between shires is drawn fixed: it is NoC-bound, not measured against clock, and spans six shire-to-shire patterns. The rings are kernels measured at 600 MHz: the matmul benchmark with tiles in L2 (fp32, fp16, int8), the same fp32 loop with every tile from DRAM, and the sparsity report's batch-1 fp32 layer with its weights in the scratchpads; the two below the ceilings sit on their bandwidth lines. The matmul rings are aifoundry2's first run and the layer aifoundry3's; the three-card check repeated each on all three cards (the tiles-in-L2 rings to 0.05%, the DRAM ring to 1%, the layer to 0.01 µs). The int8 matmul sits above the measured own-shire line (39 TOP/s at 16 OP per byte; it ran at 72) because its minions shared one tile pool (Own shire); with private tiles it ran on that line. It stays under the spec line. The tables below hold every number without the chart. A100 roofs overlays, dash-dotted, the A100's dense peaks (fp32 on the CUDA cores, fp16 and int8 on the tensor cores) and its three sustained bandwidths from the A100 table, with its ridge points on the selected precision's ceiling; the A100 lines do not move with the clock buttons.

Measured bandwidth and ridge points, as a table
Measured bandwidth and ridge points
Levelfp32fp16int8Bandwidth

Ridge points in FLOP (int8: OP) per byte at 600 MHz. Bandwidth: bytes per minion per cycle · the whole chip. For TensorSend rows, "bytes" are bytes one minion sends another.

Spec bandwidths and ridge points, as a table
Spec bandwidth and ridge points
LevelBandwidth, specSpec ridge, fp32 · fp16 · int8

The same levels at what the documented port and bank widths allow, in the same units.

The same numbers for an A100

LevelBandwidth, sustainedfp32fp16 tensorint8 tensor
How the A100 numbers were computed

Ridge points against sustained bandwidth, each ET-SoC-1 level next to its closest A100 counterpart. The A100's peaks are 19.5 TFLOP/s fp32 on the CUDA cores, 312 TFLOP/s fp16 and 624 TOP/s int8 on the tensor cores (dense). Its bandwidths are the sustained values in docs/reports/sources/2026-09-18-a100-memory-hierarchy.md (HBM2 1.40 TB/s). The ET-SoC-1's fp32 runs on its tensor unit.

The A100's tensor cores outrun its bandwidth by more, so on chip its fp16 and int8 ridge points are higher than the ET-SoC-1's. From DRAM the fp16 ridges are close (223 against 259). In int8 the ET-SoC-1 needs 2.3× more, and in IEEE fp32, which the A100 runs on its CUDA cores, about 9× more (130 against 14). With fp32 inputs rounded to TF32 on its tensor cores (156 TFLOP/s dense), the A100's HBM ridge would be about 111.

Which ridge points move with the clock

Only aifoundry2's sessions show how bandwidth follows the clock, and this section rests on its two: aifoundry3 is held at 600 MHz (a boot service sets its TDP to 0 W at every boot), while aifoundry2's governor moved it between 600 and about 800 MHz from launch to launch. Levels inside the shire run on the minion clock, because the shire cache is clocked with the minions. The L2 and scratchpad bandwidths rose exactly with the clock, so their ridge points hold at any clock. L3 and remote scratchpad reads go through the NoC, which the firmware keeps at 400 MHz. Over launches up to about 740 MHz, their bandwidth rose only 0.65× and 0.33× as fast as the minion clock. DRAM bandwidth did not change at all, so its ridge point grows in proportion to the minion clock. PCIe should behave the same way, but its rate has been timed only at 600 MHz.

fp32 ridge (FLOP/B)600 MHz800 MHz1,000 MHz
Method and caveats for the clock model

fp16 is 2× and int8 8× every entry. 600 MHz is measured on both cards; the 800 MHz (the top of the firmware's table) and 1,000 MHz (the shire design clock) columns apply the clock scaling seen on aifoundry2. Only the minion clock changes. The NoC stays at the lab cards' 400 MHz and DDR at 933 MHz; at the NoC's 500 MHz design clock the L3 and remote-scratchpad ridges would be lower. For those two, the lower end assumes their bandwidth keeps rising with the clock as it did between 600 and about 740 MHz, and the upper end assumes it stays fixed. TensorSend inside a neighbourhood or a shire runs on the minion clock, so its ridges should hold at any clock; that is inferred from the clock domain, not measured, since TensorSend has only been run at 600 MHz. Between shires it crosses the 400 MHz NoC.

Every launch, plotted against clock
Every launch's bandwidth, against its level's own 600 MHz median

One dot per launch of the memory-hierarchy report's streaming probes on aifoundry2 (18 September, both sessions), plotted at the minion clock its cycle counts imply. y = 1 is that level's own median GB/s at 600 MHz; a level whose bandwidth simply followed the clock would sit on the dashed line. L1, L2 and the own scratchpad (inside the shire, on the minion clock) track it; L3 and another shire's scratchpad (crossing the fixed-frequency NoC) rise more slowly; DRAM does not move at all.

What it takes to reach them

Matrix multiplication

How each reuse ridge is derived
Reuse to speed: how the fraction of peak grows with the block or the batch
How to read this chart

Fraction of the tensor unit's peak at 600 MHz that the measured bandwidth of each level allows (the roofline above, read along one diagonal). Blocked matmul: x is H, the harmonic mean of the C block's sides (a square block's side); the intensity is H/e per byte. L3 and another shire's scratchpad are read by one shire, whose 32 register tiles hold at most a 64 × 128 block (H = 85), so their lines stop there; DRAM assumes each byte is read once for the whole chip, with C in all 1,024 minions' registers up to H = 512 and swapped through the scratchpads (unmeasured) up to about 4,580. Streaming weights: x is the batch N; the weights give 2N/e operations per byte, and inside the own shire each op also loads a fresh activation tile, so that line flattens at N = 16. The ticks are the smallest compute-bound blocks and batches, as in the table below.

Inference that streams its weights

A layer that reads its weights once per batch of N inputs does 2N operations per weight, or 2N/e per byte (e = bytes per element). It is compute-bound when that reaches the ridge, so it needs N ≥ e × ridge / 2:

Weights come fromfp32fp16int8
Caveats on the batch table

These are the smallest batches that are compute-bound at the measured bandwidth. The own-shire row also counts the activations. A minion holds one 16×16 C tile, so each op loads a fresh activation tile as well as a weight tile. Its intensity therefore tops out at 4, 8 and 16 per byte at batch 16, and int8 never reaches its 32 OP/B own-shire ridge. The other rows assume the activations stay in the shire and each weight byte reaches the shire once. On A0 silicon, fp32 and fp16 ops of 1–4 rows with B from TenB must be padded to 5 rows (errata 1.29 type D).

A worked example: the batch-1 GEMV

At batch 1 a layer does 0.5 FLOP per byte in fp32, so it is memory-bound at every level. The sparsity report's batch-1 layer, with its weights in the scratchpads and its partial sums added by the host, ran at 1.24 TFLOP/s (6.8 µs per 1,024 × 4,096 layer: the row "Every row loaded, compute only" in that report's layer table). That is the scratchpad bandwidth times 0.5, within 1%. With the on-chip reduction the layer takes 7.3 µs, 1.15 TFLOP/s. Both times repeat to within 0.01 µs on all three cards.

Other limits that act like ridge points

Three more limits: small loads, launch overhead, the vector unit

Energy balance points

Time is one budget and energy is another. Divide each level's energy per byte by the tensor unit's energy per FLOP, both measured as card power above idle. The result is the intensity at which moving the data costs as much energy as the arithmetic. On random-normal operands held in the L1 scratchpad, the fp32 tensor unit costs 2.89 pJ per FLOP above idle [2.62–3.01] and int8 0.147 pJ per OP (energy manual §3.2, at 600 MHz; four runs on each of the three cards). By card, fp32 reads 2.96, 2.68 and 3.03 pJ per FLOP (aifoundry2, aifoundry3, aifoundry1 card 1). These are the ablation's registered values, which carry a launch-temperature offset whose size depends on how the die sensor's whole-degree reading at launch maps to the die temperature (amendment C2): aifoundry2's read from 0.6 W high to 0.2 W low, aifoundry3's 0.9–1.4 W low and aifoundry1 card 1's from 0.7 W high to 0.05 W low. So at the same die temperature aifoundry3's fp32 energy is 0.96–0.97 of aifoundry2's, not the registered 0.91 (the energy manual's catalogue, measured separately, gives 0.972), and the pooled value moves by at most 2%. The balance points below use the registered values. The data matters as much as the level: on all-ones operands fp32 costs 1.11 pJ per FLOP and on zeros 0.21, so every balance point below moves by up to 14× with the operands.

Removing the cost of keeping the cores awake

Both sides include the cost of keeping the cores awake (about 1.5–2.1 W for the whole chip; energy manual §2), which a kernel doing both pays once. Taking it out of both lowers the random-data balance points from the shire's own scratchpad to DRAM by about 9–29% (TensorSend's by more, since this cost is most of its power over idle) and changes none of the conclusions below; on zeros it is most of the tensor unit's 0.21 pJ per FLOP. On zeros the launch-temperature offset is as large as the signal: the tensor unit on zeros draws 1.85, 0.75 and 1.15 W over idle as registered, so the cards' zeros figures (0.20, 0.08 and 0.13 pJ per FLOP) cannot be compared with each other.

Energy against time: does the arithmetic or the data movement cost more once a kernel is compute-bound?
How to read this chart

One row per level, fp32 on the tensor unit: the circle is the energy balance point (byte energy ÷ FLOP energy, pooled over the three cards), the whisker every card's 99% range, the diamond the time ridge. A verdict is drawn only where every card's range clears the diamond; left of it, arithmetic costs more energy than moving the data. On zeros read the rows as a direction, not a value; the L1 row is in the table.

The energy roofline chart
The energy roofline: pJ per FLOP against arithmetic intensity, E(I) = arithmetic + data/I

All the energy figures, as a table
Where the bytes come frompJ/Bfp32 balanceint8 balancefp32 time ridge

Balance points and the time ridge are in FLOP (int8: OP) per byte. Byte energies are the energy manual's §4.2 (600 MHz, three cards, six passes each; the scratchpad rows on random data, on zeros 2.25 and 5.10 pJ/B) and §5 (TensorSend, 1 KB messages; the last row six shire-to-shire patterns). Only the own scratchpad differs between the cards beyond noise (aifoundry1 card 1 reads higher), so its row gives each card's value, aifoundry2 · aifoundry3 · aifoundry1 card 1, each balance point against its own card's FLOP energy. The int8 column pools the three cards' int8 energy (0.157, 0.131 and 0.154 pJ per OP, registered values). The L1 row sets a 32 B vector load against the vector unit's fmadd.ps (3.49 pJ per FLOP, the manual's §4.1 catalogue on random data), at that loop's 14.2 TB/s (the memory-hierarchy probe's slower loop: §4.2).

On random-normal operands, every fp32 energy balance point lies below the matching time ridge point at the measured bandwidths. So a compute-bound kernel on random-like data spends more energy on arithmetic than on moving data. The operands can change this. On all-ones operands the own scratchpad and TensorSend still lie below their time ridges, but the L2 cache, another shire's scratchpad, L3 (9.4 against 10) and DRAM (89 against 130, with the byte energy also measured on all-ones data) are not clear of theirs on every card (not established). On zeros the tensor unit adds little more power than cores that are merely awake, and the shire's own scratchpad, the L2 cache, L3 and DRAM lie well above their time ridges, even with byte energies also measured on zeros (§4.1: DRAM 434 against 130, own scratchpad 9.6 against 4.0). For zero-heavy operands, moving the bytes costs more energy than the arithmetic even when the kernel is compute-bound. Treat the zeros figures as a direction, not a value.

Checking against random-data byte energies

The §4.2 probes of the L2 cache, L3 and DRAM read whatever the buffers held, not random data. On random data a TensorLoad costs 4.21 pJ/B from the shire's own scratchpad and 129 pJ/B from DRAM (§4.1), which puts those balance points at 1.5 and 45 FLOP/B. That does not change the conclusion at the measured rates. At the spec rates 1.5 is below the shire banks' spec ridge of 2 on every card, and 45 is below DRAM's spec ridge of 72–82, but on aifoundry1 card 1 the DRAM figure's 99% range reaches 84, so that comparison is not established on every card.

Not established: comparisons too close to call

Not established

In these comparisons at least one card's 99% range reaches the ridge (the chart's tooltips give each card's range), so the page draws no verdict for them.

Caveats

Caveats in full

Reproduce

Reproduce this
python3 scripts/ridge-points.py --embed docs/reports/2026-09-18-et-soc1-ridge-points.html
python3 scripts/paste-chartkit.py docs/reports/2026-09-18-et-soc1-ridge-points.html

The script reads the raw data of the earlier reports (docs/reports/data/), the 23 September reruns of the memory probes, the on-chip communication report's embedded JSON and the energy manual's data (docs/reports/data/2026-09-23-energy-manual/manual.json). It prints the numbers behind the tables, the charts and the figures in the text, and writes the tables' and charts' data into this page. The spec constants in the script each name their source. Its inputs rebuild from committed data with the commands in docs/findings/04-artifacts.md (A19); only the raw probe and matmul runs needed a card. The second command pastes in the shared chart toolkit the charts are drawn with.

Versions. First published 24 September with energy balance points from the 18 September memory-hierarchy energies (clock free) and 1.88 pJ per FLOP on small-integer operands. Recomputed the same day from the energy manual, which moved the DRAM balance point from 79 to 42 FLOP/B and the L2 cache's below its spec ridge. Revised 25 September: the L1 row now sets a vector load against the vector unit's fmadd.ps (balance 0.15, was 0.27 against the tensor unit); the energy conclusion is stated for random-normal operands, with the reversal on zeros; the rerun check reads the passes the energy manual pools (0.3%, was 0.2%); the vector unit's L1 load rate is the energy manual catalogue loop's 23 B per minion-cycle (below about 0.69 FLOP per byte the load rate binds, was 1.6 from the memory-hierarchy probe's slower loop at 10 B); and the page gained interactive roofline, energy and reuse charts. 25 September (version 3): checked claim by claim against both cards (record: docs/reports/data/2026-09-25-claims-v3); a figure with no card named holds on both. The own scratchpad's and L2 cache's energy per byte are now per card, the int8 balance points use aifoundry2's own byte energies (DRAM 777, was 771), and an energy verdict needs both cards' 99% ranges to clear the ridge, which moves L3 at the spec rates (was "lies above") and nine other comparisons the chart drew to Not established; the all-ones DRAM comparison uses bytes measured on all-ones data (89, was 109). 26 September: the three-card check (aifoundry2, aifoundry3, aifoundry1 card 1) added to the text: the private-tile matmul (it settles the own-shire question), a lone minion's loads (floor about 42 cycles, was 45; its DRAM rate per kernel; lines in flight not settled), the short probes' ranges and the launch time per card; and the energy balance points recomputed from the energy manual's three-card data (fp32 and int8 on each card, the launch-temperature offset stated; the levels re-run on three cards, the scratchpad rows filled with random data, only the own scratchpad differing by card), which takes the L2 cache and DRAM at the spec rates, and the L2 cache and DRAM on all-ones operands, out of the verdicts (Not established, with the check's RL-X4 values). 27 September: bandwidth against the clock and the energy roofline charted (the roofline per card); the scratchpad-held block is about 4,580 on a side (was 4,500). Later that day the host link's ridge moved from the link's figure (15.75 GB/s, 624 FLOP/B in fp32) to the host-to-card DMA rate timed on three cards (the PCIe link page; about 12.5 GB/s, 786 FLOP/B), drawn solid on the roofline with the link's figure among the spec lines. 28 September: the review's cuts (the shared-tile explanation once, the energy chart's caption and table note); the note's counts given by both rules.

Sources

Sources and citations
  1. Minion VPU Specification (core-et docs, Erbium branch): §2 (8 lanes, one FMA unit and two int8 units per lane), §2.2.2 (register-file ports), §2.2.3.8 Table 6 (four int8 products per integer unit), §2.3.6.1 (tensor micro-op sequencing). ET Programmer's Reference Manual (aifoundry-org/et-man): §8.3.1 (L1 scratchpad mode), §9.1–9.4 (tensor op shapes and operands).
  2. CORE-ET Shire Cache Specification: §1 (4 banks, one 64 B line per cycle each; 512-bit NoC ports), §1.4.3 (L3 lines interleaved over all shires), §2.3–2.4 (four ports to the L3 mesh). CORE-ET Neighborhood MAS: §4.3–4.7 (neighbourhood buses, fast local network), Table 6 (shire and minion clocks). Minion DCache Description: §2.3.1 (256-bit scratchpad read port to the vector unit), §3.2 (L1 ports), §3.9.6 (TensorLoad requests in flight). CORE-ET Minion Shire Description: Table 2 (1 GHz shire and 500 MHz NoC design clocks). All in core-et.
  3. ET Preliminary Datasheet Rev 1.0 (in aifoundry-org/et-man): §1 and §7 (LPDDR4X, PCIe Gen4 x8), §2.1.1 (single issue, two harts), §7.1.1 (32 GB/s per memory-shire NoC port). Ditzel, Hot Chips 33 (2021), and Ditzel et al., IEEE Micro 42(3), 2022 (128 int8 OP per cycle per minion; 836 TOPS for six chips at 1 GHz, i.e. 139 per chip; HC33: 137 GB/s).
  4. et-platform service-processor firmware (device-bootloaders/src/ServiceProcessorBL2): common/main.c (NoC PLL at 400 MHz), driver/mem_controller.c (933 MHz DDR), include/mem_controller.h (the reported 128,000 MB/s), services/thermal_pwr_mgmt.c (minion frequency table, 300–800 MHz).
  5. P. Cawley, "Zero to matmul with the ET-SoC-1", FOSDEM 2026 (vector-unit matmul, 2.94 TFLOP/s at 650 MHz). M. Chang (marty1885), "Investigating the ET-SoC-1 NoC", 2026, also on the AI Foundry blog (about 88 GB/s from DRAM, calculated from the best rate against a single memory shire).
  6. A100: NVIDIA A100 datasheet (dense peaks, TF32 156 TFLOP/s); docs/reports/sources/2026-09-18-a100-memory-hierarchy.md (sustained shared-memory, L2 and HBM bandwidth).

← All ET-SoC-1 measurement reports