Why is the ET-SoC-1 low power?

21 September 2026, checked on three cards on 26 September · aifoundry2's card, with aifoundry3's and aifoundry1-c1's (card 1 of aifoundry1) values where the same measurement exists · A100 figures from NVIDIA's datasheet and Horace He's matmul measurements; Esperanto's design argument from Hot Chips 33 and IEEE Micro · part of the ET-SoC-1 measurement reports

Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan on aifoundry2, aifoundry3 and aifoundry1-c1 (the ablation configurations, four runs each per card). Of 20 claims tested here: 6 held, 2 corrected, 8 differ by card and 4 not confirmed, as the hub’s scoreboard counts them. The energy per multiply-add and its factor of about 19 between int8 and fp32 held on every card; the integer loop's cost and the linearity in active cores differ by card (the loop's difference lies within the reduction's launch-temperature offset, post-data note C2: section 4); the leakage split was not narrowed (the check's long idle passes, which would have re-fitted it, did not run, and its short cooling cycles on aifoundry2 could not narrow it: the DVFS report, §5). The 21 September session's numbers stay, labelled as its own (the record).

An A100 draws 330 to 400 W under a matrix multiply. This card, aifoundry2, draws 38 to 64 W at 80 °C (aifoundry3, 27 to 50 W at 56 °C), and Esperanto advertised the chip at under 20 W. Both chips are TSMC 7 nm. Esperanto's own explanation is one line, , with every term pushed down. This report measures each term on a real card and takes them away one at a time.

The short answer: the ET-SoC-1 is low power mostly because it runs at 0.52 V and 600 MHz where an A100 runs near 0.85 V and 1,160–1,410 MHz (a factor of 5 to 6 in switching power for the same capacitance), because whatever is not computing is clock-gated (an integer loop on all 1,024 cores, which Esperanto calls minions, costs 1.5 W on this card, and 0.4 and 0.7 W on aifoundry3 and aifoundry1-c1 as registered, within the launch-temperature offset; 0.0 to 1.9 W on the three cards at each launch's die temperature, section 3), and because it does 14 to 28 times fewer FLOPs per second (fp16 or fp32 here, against the A100's bf16). It is not more efficient per FLOP at dense matmul: the A100 spends 1.3 pJ per bf16 FLOP at the board, this card 3.3 pJ in fp16 and 7.0 pJ in fp32 (the other cards: the figures below). And the term Esperanto's slide leaves without a number, leakage, grows 0.65 W per degree at 80 °C and is the largest single item of the idle card: about 20 to 29 W of its 36 W at 80 °C (this card, one session; the three-card check could not narrow the split). Clock gating stops idle circuits switching, but no power gating was found on any of the three cards (a wake-up test, three probes per card), so they still leak.

Switching power, same capacitance
5–6× less
0.52 V and 600 MHz against about 0.85 V and 1,160–1,410 MHz: V² gives 2.7×, the clock 1.9–2.35×
A core spinning in an integer loop
0.4–1.5 mW
by card, 26 September: 1,024 minions add 1.5 W on this card (1.46 W on 21 September), 0.4 W on aifoundry3 and 0.7 W on aifoundry1-c1; a tensor op on zeros 1.9, 0.8 and 1.2 W; random fp32 27.2, 24.7 and 27.8 W, as registered (at the die temperature of each launch, over the two readings of post-data note C2: the loop 0.8–1.6, 1.3–1.9 and 0.0–0.8 W, zeros 1.1–1.9, 1.7–2.2 and 0.4–1.2 W, random 26.5–27.2, 25.7–26.3 and 27.1–27.9 W). The energy manual's faster addi loop costs about 1.9 mW per minion on one hart (1.9–2.0 W for 1,024 minions on the three cards) and 3.0 mW on both
Leakage at 80 °C
+0.65 W per °C
the measured slope of idle power; of the 36 W idle, 20–29 W is leakage (23 W in the best fit), a split the idle data do not pin down (this card, one session; the three-card check could not narrow it)
Energy per multiply-add, over idle
0.32 · 2.7 · 6.0 pJ
int8 · fp16 · fp32 on random data. The chip was built for the first. Held on three cards on 26 September: 0.31 · 2.7 · 5.9 pJ here, 0.26 · 2.4 · 5.4 on aifoundry3, 0.31 · 2.7 · 6.1 on aifoundry1-c1
Dense matmul, per FLOP at the board
3.3–7.0 vs 1.3 pJ
fp16–fp32 on this card (26 September: 2.6–5.5 pJ on aifoundry3, 4.1–8.5 pJ on aifoundry1-c1) against the A100's bf16 tensor cores: low power is not low energy per FLOP
Terms used on this page

The ET-SoC-1's compute cores are minions, small in-order RISC-V cores, 32 to a shire; each shire's 4 MB of SRAM is split on these cards into a 512 KB L2 cache, a 1 MB slice of the chip-wide L3 and a 2.5 MB scratchpad. A minion's tensor unit multiplies tiles with TensorFMA (one fp32 op is 4,096 multiply-adds), fed by TensorLoad, a bulk load into its L1 scratchpad. The service processor's governor sets the clock and voltage (600, 700 or 800 MHz), and the board's PMIC meters the 12 V input and three rails (minion, SRAM and the NoC, the on-chip mesh). The Horace experiment is this project's repeat of Horace He's test of matrix multiplies on predictable against random data, and its power and thermal models are used here; more in the hub's glossary.

1. The comparison

The A100 against this card, metric by metric: each bar is the ratio of the two, to the right where the A100’s value is the larger (log scale)

The comparison as a table, with sources and notes
The other A100 baseline: fp32 CUDA cores instead of tensor cores

Against the A100's fp32 CUDA-core datasheet figure instead (19.5 TFLOPS at 400 W, 20.5 pJ per FLOP, no tensor cores), this card's 7.0 pJ is about 2.9× better; the matmul efficiency report, whose faster kernel ran ±1 and ±2 operands at 57 W, puts it at 3.4×. The verdict depends on which precision the GPU is allowed. Per transistor and per square millimetre the ET-SoC-1 really is cooler, by 2.3 and 3.6 times. The rest of this report asks where that comes from.

2. Putting the terms together

Take the switching power of each chip under its matmul and split it with the equation. The A100 measured by Horace He draws 330 W under load and 88 W idle, so about 242 W switches, at about 0.85 V and 1,160–1,410 MHz: 237–288 nF of effective capacitance. This card's fp32 matmul switches 27.6 W (section 4) at 0.52 V and 600 MHz: 172 nF (on 26 September 169 nF here, 151 nF on aifoundry3 and 187 nF on aifoundry1-c1, each at its own voltage).

The factors: the A100's switching watts divided, one term at a time, down to this card's. Both ends are measured; the A100's voltage and clock under its power cap are not, so move them.
The eight terms, in a table
TermA100ET-SoC-1 (fp32)ratioHow it was established here
Switched capacitance237–288 nF172 nF1.4–1.7×power over idle ÷ V²f; close to linear in active cores on this card, 24 to 25% superlinear from 256 to 1,024 minions on the two others (section 4); 9 to 172 nF for the whole chip, depending on workload, data and precision (this card)
V²0.72 V²0.27 V²2.7×on-die voltage telemetry; in the cool-start runs on this card (two readings per pattern) the 0.62 V point costs about × per op (section 4), close to V²'s ×. The A100's voltage is an assumption: 0.75 to 0.95 V gives 2.1 to 3.3×
Clock1,160–1,410 MHz600 MHz1.9–2.35×cycle counter against wall clock here, 546 cycles per op for every data pattern; the A100's range is explained in section 1's table; the per-FLOP rows below do not depend on the clock
Switching power242 W27.6 W8.8×1.4 × 2.7 × 2.35 at 1,410 MHz; 1.7 × 2.7 × 1.9 at 1,160 MHz
Idle floor88 W36 W2.4×of the ET's, 20 to 29 W is leakage at 80 °C (23 W in the best fit; one session; the three-card check could not narrow the split); under the three-card check's runs aifoundry3 idled at 25.8 W (57 °C) and aifoundry1-c1 at 50.1 W (80 °C)
FLOPs per second257 × 10¹²9.2 × 10¹²28×bf16 tensor cores against fp32 on general vector units; in fp16 this card does 18.4 × 10¹², 14×
Capacitance switched per FLOP1.3 pF11 pF0.12×the GPU's matmul engines do 8.6 times more arithmetic per unit of switched capacitance (3.9 times against fp16, 5.0 pF)
Switching energy per FLOP0.94 pJ3.0 pJ0.31×the voltage wins back 2.7× of that 8.6×
Board energy per FLOP1.3 pJ7.0 pJ0.18×the idle floor is 57% of this card's power under load, 27% of the GPU's; in fp16 this card spends 3.3 pJ, 0.39×

So of the 8.8× lower switching power, most is the operating point: 6.3× at the A100's maximum clock (2.7× from V², 2.35× from the clock), about 5.2–5.5× at the 1,160–1,230 MHz its capped random-data run implies. The rest, 1.4–1.7×, is switching less capacitance per cycle.

Everything idle is clock-gated to nearly no switching power (an integer loop on every minion adds 1.5 W on this card; section 4 gives every card). What the card does not have is efficiency per FLOP on dense floating-point matmul: its fp32 multiply-adds switch nine times more capacitance per FLOP than the A100's bf16 tensor cores do (four times in fp16), and at 9 TFLOPS its idle floor is spread over very little work. On the workload it was designed for, int8 with sparse memory access, the arithmetic is 19 times cheaper than fp32 (18.9 to 20.5 times on three cards, section 4) and the comparison would be much closer; that was not measured against a GPU here.

3. Esperanto's equation, measured

At Hot Chips 33 Esperanto set the target as six chips on a 120 W card: under 20 W per chip, half of it for the thousand minions, "so only 10 mW per core". Their slide compares a generic x86 server core (7 W, 3 GHz, 0.85 V, 2.2 nF of switched capacitance) with the minion they needed (0.01 W, 1 GHz, 0.425 V, 0.04 nF): a reduction of 3× in frequency ("easy"), 4× in V² ("hard": circuits and SRAM) and 58× in capacitance ("very hard": architecture). The card lets each term be read off. Dividing a workload's power over idle by 1,024 minions, V² and f gives its effective switched capacitance per minion:

The capacitance tables, this card's and the three cards', and how they compare to Esperanto's target

Esperanto's target row: by P/(V²f), 10 mW at 0.425 V and 1 GHz is 0.055 nF, so the slide's 0.040 nF would leave about 3 mW for the leakage its equation includes. The measured rows are power over idle, which leaves leakage out.

The same workloads on three cards, 26 September (the claims check, four runs each; board power at the temperature each card is reduced at: aifoundry2 and aifoundry1-c1 80.9 °C, aifoundry3 55.8 °C). The idle subtracted from each run is read at its actual launch (whole-degree readings averaging 80.1, 57.4 and 80.0 °C), so the values over idle are off by the leakage slope times the gap between the die temperature at launch and the reference (post-data note C2). The reading cannot say where in its degree the die was: if each launch came on a downward step of it, as the references' launches did (die = reading + 0.96 °C), the values read 0.15 W low on aifoundry2, 1.43 W low on aifoundry3 and 0.05 W low on aifoundry1-c1; if the reading is the die temperature, 0.62 W high, 0.91 W low and 0.73 W high. That matters most for the smallest rows: at the die temperature of each launch the integer loop adds 0.8–1.6, 1.3–1.9 and 0.0–0.8 W and zeros 1.1–1.9, 1.7–2.2 and 0.4–1.2 W. The capacitance uses each card's own minion voltage under load.

The capacitance ladder: effective switched capacitance per minion for every one of the 30 ablation configurations (grey: this card, 21 September), against Esperanto's 0.040 nF target. Card marks: the twelve the three-card check re-ran on 26 September, each at its card's own voltage under load.

4. Taking the terms away one at a time

Esperanto's equation on this card: what each term contributes at a chosen voltage, clock, die temperature and number of active minions. Right: the idle law it takes leakage from, with the idle readings it was fitted to.
How the equation chart is computed

Switching is C·V²·f with each workload's capacitance measured at 0.517 V and 600 MHz (the table in section 3), scaled by the share of minions active. The fp32 matmuls are split in the proportions of the Horace experiment's flip model. The fixed part and the leakage come from the idle law, in its best-fit split (20 to 29 W of leakage at 80 °C fits the idle data as well); the extra idle measured at 0.62 V is added in proportion between the two measured voltages. Outside 0.517–0.618 V leakage was not measured, so only switching is priced there. TensorLoad streams are left out: most of their watts are in DRAM, the PHY and the SRAM rail, which the minion supply's V and f do not describe.

Activity: what is not computing costs almost nothing

Power over idle against the number of active minions, this card on 21 September and three cards on 26 September
Energy per multiply-add by precision and data. Bars: this card, 21 September; marks: each card, 26 September
The four findings behind the chart, with every number

Voltage and clock

Esperanto's modelled chip power against core voltage (Hot Chips 33), and this card's measured board power
The voltage and clock findings, with every number
Leakage and temperature, with the idle law and the thermal budget

Leakage and temperature

Memory: bandwidth traded for power, and the per-byte costs

Memory

5. Caveats

Caveats in full

6. Method and sources

Method, sources, and version history

7. Related reports