ExcelCannon

Fire an Excel workbook in — get a canonical, deterministic, LLM-ready model out. Then diff, score, and lint filled templates against it.

extractrenderdiffmigratelint

From spreadsheet to canonical model

One command reads the whole workbook — cells, formulas, styles, dropdowns, merged headers, named ranges, VBA — and writes one canonical form: JSON for machines, compact text for inference.

book.xlsx — what Excel shows

ABCD·F
1RegionQ1Q2Total
2North100120=SUM(B2:C2) 
3South9080=SUM(B3:C3) 
4East7060=SUM(B4:C4) 
5West5040=SUM(B5:C5) 
6=SUBTOTAL(109,[Q1])=SUBTOTAL(…)=SUBTOTAL(…)

bold = header row of table tblSales  ·  green = input cells (F = a Yes/No dropdown, still blank)  ·  blue = formulas  ·  r6 = the table's totals row

excelcannon
extract

canonical model — what your LLM reads

# excelcannon v5 | engine=0.13.1 | kind=data | sheets=2
= INPUT SURFACE 5 fields | 2 blank | 3 filled | 4 named | 1 unnamed
-- name Regions = Lists!$A$1:$A$4
== SHEET Form | A1:D6 | frozen r1
-- grid A1:D6
r1: Region | Q1 | Q2 | Total
r2: North | 100 | 120 | 220
r3: South | 90 | 80 | 170
r4: East | 70 | 60 | 130
r5: West | 50 | 40 | 90
r6: ~ | =SUBTOTAL(109,[Q1]) | =SUBTOTAL(109,[Q2]) | =SUBTOTAL(109,[Total])
-- fills
D2:D5 = SUM(RC[-2]:RC[-1])
-- table tblSales A1:D6 hdr totals [Region,Q1:sum,Q2:sum,Total:sum] style=TableStyleMedium2
-- validation F2:F5 list("Yes,No") blank-ok dropdown prompt="Approved?: Pick one" error="Not a choice: Choose Yes or No"
-- headers
A1:D1 [table,type,frozen] A=Region B=Q1 C=Q2 D=Total
-- inputs
Region: A2:A5 filled
Approved?: F2:F5 [list("Yes,No")] blank
-- deps
D2:D5 <- B2:C5
terminals: D2:D5 B6 C6 D6

Deterministic: the same workbook always produces byte-identical output. Sparse and compressed: a formula filled down a column is one line, not a thousand. Semantic: headers, inputs, computed cells, and dependencies are called out explicitly — the model explains how the workbook functions, not just what's in it.

Benchmark filled templates

Diff two workbooks — or a blank template against a populated one — and get category percentages plus an input-surface completion score. Built for measuring how well a template was filled.

values
90.9%
formulas
100%
structure
100%
format
100%
inputs filled
29 / 32
$ excelcannon diff expected.xlsx produced.xlsx
Sales!B3: value differs (90 -> 250)
Sales!D5: formula overwritten with literal
Overall match: 96.4%  ·  exit code 1

Share the shape, not the data

--mode schema reads a workbook and describes it instead of reproducing it. Headers, formulas, validations, tables and geometry stay; every data-region value goes. What you can hand to a model, a reviewer or a ticket when the contents are personal or confidential.

personnel.xlsx — 5 people, 22 personal values

ABCDEF
1NameEmailPhoneTeamSalaryBonus
2Dana Okonkwod.okonkwo@…+44 …Alpha68,400.00=E2*0.1
3Rhys Vantolr.vantol@…+44 …Beta71,250.00=E3*0.1
4………………

plus a hidden Directory sheet holding the approver list

excelcannon extract
--mode schema

what leaves the building (input surface and deps omitted for space)

== SHEET Roster | A1:F6 | frozen r1
-- fills
F2:F6 = RC[-1]*0.1 {fmt:#,##0.00}
-- validation D2:D6 list("Alpha,Beta") blank-ok dropdown
-- headers
A1:F1 [style,frozen] A=Name B=Email C=Phone D=Team E=Salary F=Bonus
-- schema A1:F1 -> A2:F6
A=Name (string) 5/5
B=Email (string) 5/5
D=Team (string) [list("Alpha,Beta")] 5/5
E=Salary (number) {#,##0.00} 5/5
F=Bonus (number) {#,##0.00} =RC[-1]*0.1 5/5
== SHEET Directory | A1:A3 | hidden
-- suppressed 3 cells
A1:A3 hidden 3 string

Deny by default. Text is published only for an affirmative reason — it names a column of a region the tool actually described, or it clears a set of conservative caption checks. A region the detector cannot read is withheld, not printed, and the -- suppressed block says how many cells went and what kinds they were. Every populated cell is described, published or stubbed; there is no fourth outcome. Hidden sheets are withheld whole. On a real regulatory workbook that is a schema render of 562 KB where the data render is 6.35 MB.

Reduces exposure — not an anonymizer. Schema mode is a data-minimisation tool, not a certification. Header and caption text is published by design, so a header somebody wrote a person's name into is still a header. Detection is heuristic, and a miss now costs information rather than privacy. Treat the output as a first pass and read it before sharing it.

What it captures

Every cell, sparsely

Values, formulas (A1 + R1C1), merged ranges, styles deduped to semantic properties.

Template semantics

Headers, input surface with labels and constraints, dropdowns and validations, dependency summaries.

Tables & formatting

Excel tables, conditional formats, frozen panes, hidden sheets, named ranges.

VBA source

Full macro module source extracted from .xlsm — diffable like everything else.

Binary workbooks

.xlsb reads into the same model as .xlsx — formulas included, decompiled to exactly the text Excel writes — so it diffs and lints the same way.

Lint rules

Broken refs, validation violations, overwritten formulas, error values, blank inputs — CI-ready exit codes.

Schema mode

Structure and column schemas with every data-region value suppressed, and a stub counting what was withheld.

In-process .NET SDK

The same engine as a library — P/Invoke over the native core, netstandard2.0 + net8.0, with DI registration.

Rust fast

A 132-workbook, 13.8M-cell corpus — EIOPA, EBA, ONS, filed disclosures — extracts in ~29 s total. Deterministic every time.

Install

As a command line tool

dotnet tool install -g RDLL.excelcannon

excelcannon extract book.xlsx --data --out model.json --text model.txt
excelcannon extract personnel.xlsx --mode schema --text schema.txt
excelcannon diff expected.xlsx produced.xlsx --json report.json
excelcannon migrate template_2024.xlsx template_2025.xlsx --json migration.json
excelcannon lint filled.xlsx --strict

As a library, in your process

dotnet add package RDLL.excelcannon.sdk

// the same engine, no child process — the canonical JSON is still the contract
var model = ExcelCannon.Extract("book.xlsx", ExtractMode.Schema);
Console.WriteLine(ExcelCannon.Render(model.Json, ExtractMode.Schema));

Or injected, in a host with a container — AddExcelCannon registers IExcelCannon as a singleton, idempotently, so a test can substitute it:

builder.Services.AddExcelCannon(o => { o.ExtractMode = ExtractMode.Schema; o.LintStrict = true; });

public sealed class Auditor(IExcelCannon excel)
{
    public string Summarize(string path) => excel.Render(excel.Extract(path).Json);
}

The tool ships as a .NET global tool wrapping a native Rust binary; the SDK ships the same engine as a native library it P/Invokes, targeting netstandard2.0 (.NET Framework 4.6.1+, Mono, Unity) and net8.0. Native engine for win-x64 and linux-x64 (glibc 2.17+, so RHEL 7+/Debian 10+/Ubuntu 18.04+ and the .NET container images); macOS natives are on the roadmap.