Why is the ET-SoC-1 low power?
An A100 draws 330 to 400 W under a matrix multiply. This card, aifoundry2, draws 38 to 64 W at 80 °C (aifoundry3, 27 to 50 W at 56 °C), and Esperanto advertised the chip at under 20 W. Both chips are TSMC 7 nm. Esperanto's own explanation is one line, , with every term pushed down. This report measures each term on a real card and takes them away one at a time.
The short answer: the ET-SoC-1 is low power mostly because it runs at 0.52 V and 600 MHz where an A100 runs near 0.85 V and 1,160–1,410 MHz (a factor of 5 to 6 in switching power for the same capacitance), because whatever is not computing is clock-gated (an integer loop on all 1,024 cores, which Esperanto calls minions, costs 1.5 W on this card, and 0.4 and 0.7 W on aifoundry3 and aifoundry1-c1 as registered, within the launch-temperature offset; 0.0 to 1.9 W on the three cards at each launch's die temperature, section 3), and because it does 14 to 28 times fewer FLOPs per second (fp16 or fp32 here, against the A100's bf16). It is not more efficient per FLOP at dense matmul: the A100 spends 1.3 pJ per bf16 FLOP at the board, this card 3.3 pJ in fp16 and 7.0 pJ in fp32 (the other cards: the figures below). And the term Esperanto's slide leaves without a number, leakage, grows 0.65 W per degree at 80 °C and is the largest single item of the idle card: about 20 to 29 W of its 36 W at 80 °C (this card, one session; the three-card check could not narrow the split). Clock gating stops idle circuits switching, but no power gating was found on any of the three cards (a wake-up test, three probes per card), so they still leak.
Terms used on this page
The ET-SoC-1's compute cores are minions, small in-order RISC-V cores, 32 to a shire; each shire's 4 MB of SRAM is split on these cards into a 512 KB L2 cache, a 1 MB slice of the chip-wide L3 and a 2.5 MB scratchpad. A minion's tensor unit multiplies tiles with TensorFMA (one fp32 op is 4,096 multiply-adds), fed by TensorLoad, a bulk load into its L1 scratchpad. The service processor's governor sets the clock and voltage (600, 700 or 800 MHz), and the board's PMIC meters the 12 V input and three rails (minion, SRAM and the NoC, the on-chip mesh). The Horace experiment is this project's repeat of Horace He's test of matrix multiplies on predictable against random data, and its power and thermal models are used here; more in the hub's glossary.
1. The comparison
The comparison as a table, with sources and notes
The other A100 baseline: fp32 CUDA cores instead of tensor cores
Against the A100's fp32 CUDA-core datasheet figure instead (19.5 TFLOPS at 400 W, 20.5 pJ per FLOP, no tensor cores), this card's 7.0 pJ is about 2.9× better; the matmul efficiency report, whose faster kernel ran ±1 and ±2 operands at 57 W, puts it at 3.4×. The verdict depends on which precision the GPU is allowed. Per transistor and per square millimetre the ET-SoC-1 really is cooler, by 2.3 and 3.6 times. The rest of this report asks where that comes from.
2. Putting the terms together
Take the switching power of each chip under its matmul and split it with the equation. The A100 measured by Horace He draws 330 W under load and 88 W idle, so about 242 W switches, at about 0.85 V and 1,160–1,410 MHz: 237–288 nF of effective capacitance. This card's fp32 matmul switches 27.6 W (section 4) at 0.52 V and 600 MHz: 172 nF (on 26 September 169 nF here, 151 nF on aifoundry3 and 187 nF on aifoundry1-c1, each at its own voltage).
The eight terms, in a table
| Term | A100 | ET-SoC-1 (fp32) | ratio | How it was established here |
|---|---|---|---|---|
| Switched capacitance | 237–288 nF | 172 nF | 1.4–1.7× | power over idle ÷ V²f; close to linear in active cores on this card, 24 to 25% superlinear from 256 to 1,024 minions on the two others (section 4); 9 to 172 nF for the whole chip, depending on workload, data and precision (this card) |
| V² | 0.72 V² | 0.27 V² | 2.7× | on-die voltage telemetry; in the cool-start runs on this card (two readings per pattern) the 0.62 V point costs about × per op (section 4), close to V²'s ×. The A100's voltage is an assumption: 0.75 to 0.95 V gives 2.1 to 3.3× |
| Clock | 1,160–1,410 MHz | 600 MHz | 1.9–2.35× | cycle counter against wall clock here, 546 cycles per op for every data pattern; the A100's range is explained in section 1's table; the per-FLOP rows below do not depend on the clock |
| Switching power | 242 W | 27.6 W | 8.8× | 1.4 × 2.7 × 2.35 at 1,410 MHz; 1.7 × 2.7 × 1.9 at 1,160 MHz |
| Idle floor | 88 W | 36 W | 2.4× | of the ET's, 20 to 29 W is leakage at 80 °C (23 W in the best fit; one session; the three-card check could not narrow the split); under the three-card check's runs aifoundry3 idled at 25.8 W (57 °C) and aifoundry1-c1 at 50.1 W (80 °C) |
| FLOPs per second | 257 × 10¹² | 9.2 × 10¹² | 28× | bf16 tensor cores against fp32 on general vector units; in fp16 this card does 18.4 × 10¹², 14× |
| Capacitance switched per FLOP | 1.3 pF | 11 pF | 0.12× | the GPU's matmul engines do 8.6 times more arithmetic per unit of switched capacitance (3.9 times against fp16, 5.0 pF) |
| Switching energy per FLOP | 0.94 pJ | 3.0 pJ | 0.31× | the voltage wins back 2.7× of that 8.6× |
| Board energy per FLOP | 1.3 pJ | 7.0 pJ | 0.18× | the idle floor is 57% of this card's power under load, 27% of the GPU's; in fp16 this card spends 3.3 pJ, 0.39× |
So of the 8.8× lower switching power, most is the operating point: 6.3× at the A100's maximum clock (2.7× from V², 2.35× from the clock), about 5.2–5.5× at the 1,160–1,230 MHz its capped random-data run implies. The rest, 1.4–1.7×, is switching less capacitance per cycle.
Everything idle is clock-gated to nearly no switching power (an integer loop on every minion adds 1.5 W on this card; section 4 gives every card). What the card does not have is efficiency per FLOP on dense floating-point matmul: its fp32 multiply-adds switch nine times more capacitance per FLOP than the A100's bf16 tensor cores do (four times in fp16), and at 9 TFLOPS its idle floor is spread over very little work. On the workload it was designed for, int8 with sparse memory access, the arithmetic is 19 times cheaper than fp32 (18.9 to 20.5 times on three cards, section 4) and the comparison would be much closer; that was not measured against a GPU here.
3. Esperanto's equation, measured
At Hot Chips 33 Esperanto set the target as six chips on a 120 W card: under 20 W per chip, half of it for the thousand minions, "so only 10 mW per core". Their slide compares a generic x86 server core (7 W, 3 GHz, 0.85 V, 2.2 nF of switched capacitance) with the minion they needed (0.01 W, 1 GHz, 0.425 V, 0.04 nF): a reduction of 3× in frequency ("easy"), 4× in V² ("hard": circuits and SRAM) and 58× in capacitance ("very hard": architecture). The card lets each term be read off. Dividing a workload's power over idle by 1,024 minions, V² and f gives its effective switched capacitance per minion:
The capacitance tables, this card's and the three cards', and how they compare to Esperanto's target
Esperanto's target row: by P/(V²f), 10 mW at 0.425 V and 1 GHz is 0.055 nF, so the slide's 0.040 nF would leave about 3 mW for the leakage its equation includes. The measured rows are power over idle, which leaves leakage out.
The same workloads on three cards, 26 September (the claims check, four runs each; board power at the temperature each card is reduced at: aifoundry2 and aifoundry1-c1 80.9 °C, aifoundry3 55.8 °C). The idle subtracted from each run is read at its actual launch (whole-degree readings averaging 80.1, 57.4 and 80.0 °C), so the values over idle are off by the leakage slope times the gap between the die temperature at launch and the reference (post-data note C2). The reading cannot say where in its degree the die was: if each launch came on a downward step of it, as the references' launches did (die = reading + 0.96 °C), the values read 0.15 W low on aifoundry2, 1.43 W low on aifoundry3 and 0.05 W low on aifoundry1-c1; if the reading is the die temperature, 0.62 W high, 0.91 W low and 0.73 W high. That matters most for the smallest rows: at the die temperature of each launch the integer loop adds 0.8–1.6, 1.3–1.9 and 0.0–0.8 W and zeros 1.1–1.9, 1.7–2.2 and 0.4–1.2 W. The capacitance uses each card's own minion voltage under load.
- Capacitance: the architecture term holds. An integer loop switches 0.009 nF per minion, a fully gated tensor op 0.012 nF, int8 multiply-adds on random data 0.061 nF (this card, 21 September; on the three cards of 26 September the loop 0.002 to 0.009 nF, int8 on random data 0.050 to 0.064 nF and fp32 on random data 0.148 to 0.183 nF). The slide's 0.04 nF sits in the middle of the chip's intended workloads. Only floating-point multiply-adds on random data, which the chip was not designed around, reach 0.15 to 0.17 nF.
- 10 mW per core: the switching part fits for int8, even at this card's higher voltage: 4 mW over idle on constant data and 10 mW on random data (3.0 to 4.3 and 8.1 to 9.7 mW on the three cards of 26 September). But Esperanto's 10 mW was the whole budget, leakage included, and this card's minion rail alone reads 22 W under int8 random data at the end of a 7 s run (about 83 °C), 21.7 mW per minion. fp32 on random data switches 27 mW per core (24.1 to 27.2 mW on the three cards).
- Voltage and clock: this card is not at the advertised point. Its lowest operating point is 0.52 V at 600 MHz. While the die reads 65 °C or less (whole degrees) and board power is under 65 W, this card's TDP setting lets the firmware step it up through 0.57 V at 700 MHz to 0.62 V at 800 MHz (the DVFS loop, and the leakage, which also says why the other two cards stay at 600 MHz). Esperanto's "about 0.4 V, 20 W" is a different operating point from anything this card's firmware uses.
4. Taking the terms away one at a time
How the equation chart is computed
Switching is C·V²·f with each workload's capacitance measured at 0.517 V and 600 MHz (the table in section 3), scaled by the share of minions active. The fp32 matmuls are split in the proportions of the Horace experiment's flip model. The fixed part and the leakage come from the idle law, in its best-fit split (20 to 29 W of leakage at 80 °C fits the idle data as well); the extra idle measured at 0.62 V is added in proportion between the two measured voltages. Outside 0.517–0.618 V leakage was not measured, so only switching is priced there. TensorLoad streams are left out: most of their watts are in DRAM, the PHY and the SRAM rail, which the minion supply's V and f do not describe.
Activity: what is not computing costs almost nothing
The four findings behind the chart, with every number
- The general-purpose side is nearly free. 1,024 minions spinning in an integer loop add 1.46 W to the idle card: 1.4 mW per core, 8 pJ per instruction (this card, 21 September, two runs). The three-card check measured 1.5 W [0.6, 2.4] on this card, 0.4 W [−0.1, 0.9] on aifoundry3 and 0.7 W [−0.7, 2.1] on aifoundry1-c1 (four runs each): only this card's excludes zero. Those are the registered values, which carry each card's launch-temperature offset (post-data note C2, section 3's second table). At the die temperature of each launch the loop adds 0.8–1.6 W on this card, 1.3–1.9 W on aifoundry3 and 0.0–0.8 W on aifoundry1-c1 over the note's two readings of the whole-degree launch temperature; with the intervals the registered item computes, only this card's excludes zero, and only if the launches came on a downward step of the reading (1.6 W [0.6, 2.5]; 0.8 W [−0.1, 1.7] if the reading is the die temperature). That loop, four adds and a branch, issues about 0.3 instructions per cycle on hart 0; the energy manual's tighter addi loop draws about 1.9 mW per minion on one hart (1.87 mW on this card; 1.9–2.0 W for 1,024 minions on each of the three cards) and 3.0 mW on both, and with both harts a nop or a fence costs 4.4–5.1 pJ per issue slot. Esperanto's claim that RISC-V compatibility costs little is borne out. A fully gated TensorFMA (zeros) costs 1.9 W for all cores on this card (1.9, 0.8 and 1.2 W on the three cards on 26 September; 1.1–1.9, 1.7–2.2 and 0.4–1.2 W at the die temperature of each launch); presumably the integer pipeline sleeps during a tensor instruction.
- Power is close to linear in active cores on this card: 25.6 mW per minion at 256 and 512 active, 26.2 at 768 and 27.0 at 1,024; a line through zero fits 26.5 mW per minion (21 September). The three-card check of 26 September found it superlinear: per active minion, 1,024 minions switch 9% more than 256 on this card (ratio 1.09 [0.75, 1.59]) and 24 and 25% more on aifoundry3 and aifoundry1-c1 (1.24 [1.10, 1.39] and 1.25 [1.01, 1.53]), outside the ±10% registered as linear. In the fp32 matmul idle cores cost nothing measurable: on every card a line through 256, 512 and 1,024 minions meets zero minions 0.9 to 2.1 W below the card's idle, so there is no floor above it. The integer loop is the exception: 0.54 W on 256 minions and 1.46 W on 1,024 (this card, 21 September, two runs each) put its line through zero minions at about 0.24 W.
- Gating follows the data. At the same FLOPs the fp32 matmul adds 1.9 W on zeros, 10.6 W on ones and 27.6 W on random values (1.9, 9.7 and 24.9 W on aifoundry3, 22 September). A multiply-add whose operand is zero gets no valid bit and clocks no register; a constant clocks registers but toggles no data. The Horace experiment predicts these from the RTL to 0.5 W rms in one session, 0.9 to 1.0 W on the three-card check's patterns (section 3), and 14 structured matrices priced before they ran to 0.9 W rms (section 9).
- Precision is the biggest lever inside the chip. On random data an int8 multiply-add costs 0.32 pJ over idle, fp16 2.7 pJ and fp32 6.0 pJ: a factor of 19 between int8, the type the chip was built for, and fp32. It held on all three cards on 26 September: 0.31, 2.7 and 5.9 pJ on this card, 0.26, 2.4 and 5.4 pJ on aifoundry3 and 0.31, 2.7 and 6.1 pJ on aifoundry1-c1, with fp32 over int8 18.9 [18.3, 19.5], 20.5 [19.0, 22.0] and 19.6 [17.9, 21.5] times. The int8 unit also runs 6.9 times more multiply-adds per second (51 per cycle per minion against 7.5).
Voltage and clock
The voltage and clock findings, with every number
- The card has three operating points (600, 700 and 800 MHz); on random data the two ends differ about as CV²f says. This is aifoundry2 alone, in the seven cool-start runs of one morning (21 September; see Caveats). From a cool die the firmware's clock governor runs kernels at 800 MHz and 0.62 V; past 65 °C, or 65 W at the board, it steps back down to 600 MHz and 0.52 V. Random fp32, at nearly the same die temperature ( °C), draws W over idle at 800 MHz against W at 600 MHz: × the switching power for 1.33× the clock, where V²f predicts ×. Ones give × ( against W; their two runs' 800 MHz readings agree to 0.2 W). Zeros cannot test it: their two runs switch 4.0 and 4.5 W at 800 MHz, and the 600 MHz figure, W, comes from the 80 °C session, since from a cool die zeros barely left 800 MHz. Energy per operation rises about × for 33% more speed.
- Idle pays for voltage too: W at 0.62 V (two half-second stretches) against W at 0.52 V, both at °C, a quarter more for a tenth of a volt; how much of that is leakage and how much the always-running clocks was not measured.
- Extrapolated to a GPU's operating point, the same silicon is a GPU-class part. Random fp32 switches 27.6 W here; at 0.85 V and 1,410 MHz the same capacitance would switch 175 W, before leakage, which also grows with voltage. Esperanto's own model is in the same range: about 230 W at 0.85 V, 164 W at 0.75 V. This card's measured board power at 0.52 V is 47 W on ones, within the 44 to 49 W the curve gives there (depending on how its published points are interpolated), and 64 W on random fp32, above it; the curve is chip power at its own clock and workload.
Leakage and temperature, with the idle law and the thermal budget
Leakage and temperature
- Idle board power follows a law fitted to this card's idle readings: 0.65 W more per degree at 80 °C, with 20 to 29 W of the idle at 80 °C in leakage. The law, its checks on three cards and its split are in the DVFS report, §5 and the energy manual, §1.
- This is the term that keeps the measured card from Esperanto's headline. The minion, SRAM and NoC rails alone read 22 W with every multiply-add gated, at 80 °C (13.9 W on aifoundry3 at 56 °C). A cooler die and the 0.4 V operating point would both cut it: leakage falls exponentially with temperature and steeply with voltage.
- It also makes the card thermally fragile: leakage feeds its own heat back, which leaves a budget of about 3 W of sustained switching on this card's cooling (the Horace experiment, section 8).
Memory: bandwidth traded for power, and the per-byte costs
Memory
- Moving data costs power too: on this card a TensorLoad stream from LPDDR4x at GB/s ( of the 119 GB/s these cards' DDR clock allows) adds W on a never-written buffer (end to end: DRAM, PHY, controller, mesh and cache fills), and one from the shire's own L2 at TB/s W (two 7 s runs each at 80 °C). Per byte, use the energy manual, §4, which has the values with bars on three cards: DRAM 114.6 pJ/B [89.0–141.3] at 600 MHz on buffers whose contents are not set (95 on zeros and 133 on random data by tensor load), the L2 2.6 pJ/B on aifoundry2, 2.8 on aifoundry3 and 3.9 on aifoundry1-c1.
- An A100's HBM moves 1,555 to 2,039 GB/s at peak, 13 to 17 times the 119 GB/s these cards' DDR clock allows (11 to 15 times the LPDDR4x datasheet maximum of 137 GB/s in the section 1 table). Esperanto's choice trades bandwidth per chip for power, cost and capacity, and gets bandwidth back by using six chips per card.
5. Caveats
Caveats in full
- No A100 was measured here. Its numbers are the datasheet's and Horace He's (one GPU, bf16 8192³ matmul, a 330 W limit). Its core voltage is not published; 0.85 V is an assumption, and the table gives the range. Nor is its clock under the 330 W cap: 1,410 MHz is an upper bound.
- Everything is board power, LPDDR4x, regulators and PCIe included (Limits of observability, section 4.2).
- The 800 MHz figures come from seven short runs on one morning in which the governor changed the clock within seconds.
Each is the highest board reading before the governor stepped down; random data held 800 MHz for at most 0.3 s, and the two
of its three runs that caught a reading there agree to 0.3 W. Each is a single highest reading over an idle level that moves
with the die temperature (the 600 MHz idle alone by about W per degree at 64–66 °C), so random
data's and ones' above V²f are as close as these runs can test it.
They deserve a dedicated run from a cool die (
tools/ettelem/run_vf_cold.shis written for it), which only hours idle in a cool room give (next item). They are also one card's: aifoundry3 stays at 600 MHz (a boot service sets its TDP to 0 W), and aifoundry1-c1's governor does not raise its clock, even busy on a die of 64 °C or less (the DVFS report, section 3). - The card sits in a desktop chassis: left alone it read 62 °C after the night of 20–21 September and 73 °C after 20.6 hours on a warmer 22 September (one reading each), and every run here launched at 80 °C. In a server airflow it would presumably run cooler and leak less.
- Mostly one card. The three-card check of 26 September repeated the ablation configurations on aifoundry2, aifoundry3 and aifoundry1-c1, four runs each. aifoundry1-c1, launched at this card's 80 °C, switched 1% more than it on the fp32 operand patterns (least squares over eight patterns); aifoundry3, launched 23 °C cooler, switched 10% less. Most of that 10% is the reduction, not the card (post-data note C2): at the same die temperature aifoundry3 switches 3–4% less (0.958–0.969) and aifoundry1-c1 0.4% more (1.004), as the energy manual's independent catalogue ratio, 0.972, found. A test at two launch temperatures on each card could not say whether the rest is the card or its temperature (the Horace experiment, section 10; the DVFS report, section 6), so the capacitances here are this card's at 80 °C; where the others ran the same measurement their values are given beside them. Every card's registered switching values carry that launch-temperature offset; section 3's second table gives its size, which matters most for the smallest values, the integer loop and zeros.
6. Method and sources
Method, sources, and version history
- Evidence: this page rests on aifoundry2's card. The claims check of 26 September repeated the ablation configurations (precisions, data, active minions, the integer loop) on aifoundry2, aifoundry3 and aifoundry1-c1, four runs each, and its values are given beside the claims it tested and in section 3's second table; the 800 MHz points and the long runs are aifoundry2 only. Counts of runs are given where they are fewer than three. The tests and intervals are in the claims-check record.
- Every configuration ran twice for 7 s under the strict start of the Horace experiment (heat to 84 °C if below, idle to the 81→80 °C step, launch), in shuffled order, with 10 Hz telemetry. Power is the mean over seconds 1 to 3 moved to the launch temperature with the measured leakage slope. Repeats agree to 0.05 W for most configurations (this session; in the check's four runs per configuration and card, half the range of a configuration's runs was 0.08 to 0.21 W in the median over the three cards, and up to 0.5 W). These are the ablation session's values (random fp32 W at the launch temperature). The Horace experiment's strict session the same day quotes 63.4 W (+27.1 W over idle): the two sessions agree to 0.2 W, and the pages quote 63.9 and 63.4 W because they move power to the launch temperature with different leakage slopes, 0.68 W per degree here and 0.81 there.
- Commands:
tools/ettelem/run_ablation.sh build/ablation1 tools/ettelem/ablation.cfg 2 7, thentools/ettelem/analyze_ablation.py.tools/ettelem/finish_horace.shrebuilds this page and the Horace experiment from the data kept in the repository. Measurements usedtools/ettelem,workloads/sparsityand the RTL flip counts ofrtl-sim/fma_toggle. - Data, all in
docs/reports/data/2026-09-21-horace-aifoundry2/:ablation.jsonandablation.txt(every configuration's runs),vf.json(the two operating points, fromtools/ettelem/build_vf.py) andmodel.json(the idle law and the flip model). Notes and quotes:docs/research/why-low-power.md, written on 20–21 September; where its measured numbers differ, this page supersedes them. - Firmware: the governor's source was read at et-platform
353f20e; the card's own trace strings match an older build (before et-platform commit60b40c10f, 24 September 2024). The cards' own build (BL2 0.20.0 of release 1.3.1, et-platformffca4cbb4) was read on 27 September: its thermal response is a blocking loop of steps about 0.4 s apart that also acts on an idle card, and a climb goes to the top point in one call (the DVFS report). The three operating points show in this card's telemetry. - D. Ditzel et al., "Accelerating ML Recommendation with over a Thousand RISC-V/Tensor Processors on Esperanto's ET-SoC-1 Chip", Hot Chips 33 slides, 2021, and IEEE Micro 42(3), 2022.
- NVIDIA A100 datasheet. H. He, "Strangely, Matrix Multiplications on GPUs Run Faster When Given 'Predictable' Data!", 2024.
- Versions. 21 September 2026: first published. 24 September: corrected after review (the card's three operating points, the A100's clock under its cap, the fp16 comparison, the thermal budget's dependence on the day). 25 September: the V²f check restated on random data; then version 3, every claim checked against both cards (the leakage headline became the measured slope with the split as a range; the energies per byte are the energy manual's). 26 September (version 4): the three-card check's values beside the tested claims; the linearity in active cores and the integer loop per card. 27 September: the review's fixes ("brackets" corrected; the integer loop's floor) and a capacitance ladder. 28 September: aifoundry1-c1's governor never raised the clock; later, the review's cuts. What each version changed in full: this page's history in the repository.
7. Related reports
- The Horace experiment — the flip model and the thermal model used here, and the strict start every run followed.
- The DVFS loop and its leakage — the three operating points and when the governor moves between them, the idle law checked session by session on three cards, and the card whose clock cannot rise.
- The energy manual — per-event costs with bars from repeated passes on three cards, including the awake core and the bytes read from each memory level.
- Matmul efficiency — the same tensor unit in a faster matmul loop, set against the A100's fp32 datasheet figure rather than its tensor cores.
- Limits of observability, section 4 — what the three rails do not meter, and where the unmetered watts go.
- Power and temperature — how often each power and temperature reading updates, and how idle power rises with die temperature.
- Memory hierarchy — latency, bandwidth and energy per byte of each memory level, first on this card and since on three cards at 600 MHz (its energies are the energy manual's, §4).