ExcelCannon
Fire an Excel workbook in — get a canonical, deterministic, LLM-ready model out. Then diff, score, and lint filled templates against it.
extractrenderdiffmigratelint
From spreadsheet to canonical model
One command reads the whole workbook — cells, formulas, styles, dropdowns, merged headers, named ranges, VBA — and writes one canonical form: JSON for machines, compact text for inference.
book.xlsx — what Excel shows
| A | B | C | D | · | F | |
|---|---|---|---|---|---|---|
| 1 | Region | Q1 | Q2 | Total | ||
| 2 | North | 100 | 120 | =SUM(B2:C2) | ||
| 3 | South | 90 | 80 | =SUM(B3:C3) | ||
| 4 | East | 70 | 60 | =SUM(B4:C4) | ||
| 5 | West | 50 | 40 | =SUM(B5:C5) | ||
| 6 | =SUBTOTAL(109,[Q1]) | =SUBTOTAL(…) | =SUBTOTAL(…) |
bold = header row of table tblSales · green = input cells (F = a Yes/No dropdown, still blank) · blue = formulas · r6 = the table's totals row
extract
canonical model — what your LLM reads
# excelcannon v5 | engine=0.13.1 | kind=data | sheets=2 = INPUT SURFACE 5 fields | 2 blank | 3 filled | 4 named | 1 unnamed -- name Regions = Lists!$A$1:$A$4 == SHEET Form | A1:D6 | frozen r1 -- grid A1:D6 r1: Region | Q1 | Q2 | Total r2: North | 100 | 120 | 220 r3: South | 90 | 80 | 170 r4: East | 70 | 60 | 130 r5: West | 50 | 40 | 90 r6: ~ | =SUBTOTAL(109,[Q1]) | =SUBTOTAL(109,[Q2]) | =SUBTOTAL(109,[Total]) -- fills D2:D5 = SUM(RC[-2]:RC[-1]) -- table tblSales A1:D6 hdr totals [Region,Q1:sum,Q2:sum,Total:sum] style=TableStyleMedium2 -- validation F2:F5 list("Yes,No") blank-ok dropdown prompt="Approved?: Pick one" error="Not a choice: Choose Yes or No" -- headers A1:D1 [table,type,frozen] A=Region B=Q1 C=Q2 D=Total -- inputs Region: A2:A5 filled Approved?: F2:F5 [list("Yes,No")] blank -- deps D2:D5 <- B2:C5 terminals: D2:D5 B6 C6 D6
Deterministic: the same workbook always produces byte-identical output. Sparse and compressed: a formula filled down a column is one line, not a thousand. Semantic: headers, inputs, computed cells, and dependencies are called out explicitly — the model explains how the workbook functions, not just what's in it.
Benchmark filled templates
Diff two workbooks — or a blank template against a populated one — and get category percentages plus an input-surface completion score. Built for measuring how well a template was filled.
$ excelcannon diff expected.xlsx produced.xlsx Sales!B3: value differs (90 -> 250) Sales!D5: formula overwritten with literal Overall match: 96.4% · exit code 1
Share the shape, not the data
--mode schema reads a workbook and describes it instead of reproducing it. Headers, formulas, validations, tables and geometry stay; every data-region value goes. What you can hand to a model, a reviewer or a ticket when the contents are personal or confidential.
personnel.xlsx — 5 people, 22 personal values
| A | B | C | D | E | F | |
|---|---|---|---|---|---|---|
| 1 | Name | Phone | Team | Salary | Bonus | |
| 2 | Dana Okonkwo | d.okonkwo@… | +44 … | Alpha | 68,400.00 | =E2*0.1 |
| 3 | Rhys Vantol | r.vantol@… | +44 … | Beta | 71,250.00 | =E3*0.1 |
| 4 | … | … | … | … | … | … |
plus a hidden Directory sheet holding the approver list
--mode schema
what leaves the building (input surface and deps omitted for space)
== SHEET Roster | A1:F6 | frozen r1 -- fills F2:F6 = RC[-1]*0.1 {fmt:#,##0.00} -- validation D2:D6 list("Alpha,Beta") blank-ok dropdown -- headers A1:F1 [style,frozen] A=Name B=Email C=Phone D=Team E=Salary F=Bonus -- schema A1:F1 -> A2:F6 A=Name (string) 5/5 B=Email (string) 5/5 D=Team (string) [list("Alpha,Beta")] 5/5 E=Salary (number) {#,##0.00} 5/5 F=Bonus (number) {#,##0.00} =RC[-1]*0.1 5/5 == SHEET Directory | A1:A3 | hidden -- suppressed 3 cells A1:A3 hidden 3 string
Deny by default. Text is published only for an affirmative reason — it names a column of a region the tool actually described, or it clears a set of conservative caption checks. A region the detector cannot read is withheld, not printed, and the -- suppressed block says how many cells went and what kinds they were. Every populated cell is described, published or stubbed; there is no fourth outcome. Hidden sheets are withheld whole. On a real regulatory workbook that is a schema render of 562 KB where the data render is 6.35 MB.
Reduces exposure — not an anonymizer. Schema mode is a data-minimisation tool, not a certification. Header and caption text is published by design, so a header somebody wrote a person's name into is still a header. Detection is heuristic, and a miss now costs information rather than privacy. Treat the output as a first pass and read it before sharing it.
What it captures
Values, formulas (A1 + R1C1), merged ranges, styles deduped to semantic properties.
Headers, input surface with labels and constraints, dropdowns and validations, dependency summaries.
Excel tables, conditional formats, frozen panes, hidden sheets, named ranges.
Full macro module source extracted from .xlsm — diffable like everything else.
.xlsb reads into the same model as .xlsx — formulas included, decompiled to exactly the text Excel writes — so it diffs and lints the same way.
Broken refs, validation violations, overwritten formulas, error values, blank inputs — CI-ready exit codes.
Structure and column schemas with every data-region value suppressed, and a stub counting what was withheld.
The same engine as a library — P/Invoke over the native core, netstandard2.0 + net8.0, with DI registration.
A 132-workbook, 13.8M-cell corpus — EIOPA, EBA, ONS, filed disclosures — extracts in ~29 s total. Deterministic every time.
Install
As a command line tool
dotnet tool install -g RDLL.excelcannon excelcannon extract book.xlsx --data --out model.json --text model.txt excelcannon extract personnel.xlsx --mode schema --text schema.txt excelcannon diff expected.xlsx produced.xlsx --json report.json excelcannon migrate template_2024.xlsx template_2025.xlsx --json migration.json excelcannon lint filled.xlsx --strict
As a library, in your process
dotnet add package RDLL.excelcannon.sdk
// the same engine, no child process — the canonical JSON is still the contract
var model = ExcelCannon.Extract("book.xlsx", ExtractMode.Schema);
Console.WriteLine(ExcelCannon.Render(model.Json, ExtractMode.Schema));
Or injected, in a host with a container — AddExcelCannon registers IExcelCannon as a singleton, idempotently, so a test can substitute it:
builder.Services.AddExcelCannon(o => { o.ExtractMode = ExtractMode.Schema; o.LintStrict = true; });
public sealed class Auditor(IExcelCannon excel)
{
public string Summarize(string path) => excel.Render(excel.Extract(path).Json);
}
The tool ships as a .NET global tool wrapping a native Rust binary; the SDK ships the same engine as a native library it P/Invokes, targeting netstandard2.0 (.NET Framework 4.6.1+, Mono, Unity) and net8.0. Native engine for win-x64 and linux-x64 (glibc 2.17+, so RHEL 7+/Debian 10+/Ubuntu 18.04+ and the .NET container images); macOS natives are on the roadmap.