An MLIR-based domain-specific compiler for synthetic aperture radar imaging
Write SAR imaging pipelines in Python at the granularity of whole stages — FFTs, phase multiplies, Stolt interpolation, corner turns — and compile them to native CPU code or synthesizable FPGA designs.
San Francisco Bay: 16384 x 16384 raw ALOS-1 echoes focused by a single compiled omega-K kernel in 3.7 seconds.
- Three complete imaging algorithms. omega-K (WKA), Range-Doppler (RDA) and Chirp Scaling (CSA) chains, each a single compiled kernel, validated numerically against NumPy references, cross-checked against each other on point targets, and demonstrated on real satellite data (examples/).
- Multi-backend by construction. A
cpubackend executes kernels natively (linalg fusion, OpenMP,libsar_runtimeFFT); ascalehlsbackend emits Vitis HLS C++ through ScaleHLS-HIDA. Every SAR operation has an HLS lowering — FFTs as bit-reversal-free Stockham affine loops, interpolation as straight-line masked gathers — so complete imaging chains emit as single FPGA designs. New targets plug in as backend packages without core changes. - A real dialect, not a wrapper. Domain invariants (power-of-two FFT
sizes, frequency-axis consistency, precision rules) are enforced by MLIR
verifiers; bit-exact canonicalization folds; orthogonal primitives
(
sar.interp1dunderlies both Stolt remapping and RCMC). - Tested like a compiler. FileCheck suites for every pass, per-op numerics against NumPy, end-to-end algorithm equivalence, memory-leak regression, and a differential fuzzer comparing random DSL programs against a NumPy oracle.
import sar
N = 512
@sar.jit
def range_compress(raw: sar.c64[N, N], ref: sar.c64[N, N]) -> sar.c64[N, N]:
spectrum = sar.fftshift(sar.fft(raw, dim=1), dim=1)
return sar.ifft(sar.ifftshift(spectrum * ref, dim=1), dim=1)
image = range_compress(raw_np, ref_np) # JIT on CPU
design = range_compress.compile(backend="scalehls") # or emit HLS C++
print(design.cpp_path)Calling a kernel traces the Python function into a small MLIR dialect of whole-array SAR operations; the selected backend then runs a pipeline of compilation stages over it (cached on disk, keyed by content):
flowchart TB
K["@sar.jit kernel<br/><sub>Python tracing, shape/dtype checks</sub>"]
IR["sar dialect module<br/><sub>textual MLIR, verified by sar-opt</sub>"]
K --> IR
IR -- "sar-to-llvm-pipeline" --> CPU1
IR -- "sar-to-linalg-pipeline<br/><sub>float element-wise</sub>" --> H1
IR -- "sar-to-affine-pipeline<br/><sub>complex / FFT</sub>" --> H2
subgraph CPU["cpu backend · native execution"]
direction TB
CPU1["linalg fusion · OpenMP parallel loops<br/>FFT and interpolation → libsar_runtime"]
CPU2["mlir-translate → clang -O3"]
CPU3["kernel.so<br/><sub>ctypes launcher, memref ABI</sub>"]
CPU1 --> CPU2 --> CPU3
end
subgraph HLS["scalehls backend · FPGA emission"]
direction TB
H1["HIDA PyTorch flow<br/><sub>dataflow, tiling</sub>"]
H2["decomplexify → Stockham FFT affine loops<br/>→ HIDA C++ flow"]
H3["kernel.hls.cpp<br/><sub>Vitis HLS C++ with directives</sub>"]
H1 --> H3
H2 --> H3
end
The design decisions behind this shape — and their trade-offs — are laid out in docs/architecture.md. The short version:
| Decision | Why |
|---|---|
| Ops on builtin tensors (like TOSA/StableHLO) | no custom type conversions; all upstream MLIR machinery just works |
| Textual IR as the frontend boundary | pure-Python frontend, zero LLVM API coupling; sar-opt verifiers remain authoritative |
| Runtime library for CPU FFT/interpolation | the vendor-library pattern (cf. cuFFT): pipeline structure is the compiler's job, leaf transforms are not |
| Split-complex + Stockham FFT for HLS | no bit-reversal means fully affine accesses — synthesizable, and validated against NumPy on the CPU |
| Interpolation as straight-line masked gathers | clamped indices + select-masked taps instead of control flow, keeping HLS loop pipelining applicable |
Full omega-K imaging chain, 240-core x86-64 server (benchmarks/):
| Scene | SAR-DSL (cpu) | NumPy reference | Speedup |
|---|---|---|---|
| 1024 x 1024 synthetic | 0.23 s | 0.46 s | 2.0x |
| 4096 x 4096 synthetic | 0.74 s | 6.34 s | 8.6x |
| 16384 x 16384 ALOS-1 | 3.7 s | ~15 min* | ~240x |
* extrapolated; the reference Stolt loop is impractical at this size.
Requirements: cmake >= 3.20, ninja, a C++17 host compiler, Python >= 3.10 with numpy (matplotlib for the examples).
git clone <repo> && cd sar-dsl
git submodule update --init externals/llvm-project
pip install numpy matplotlib pytest # matplotlib/pytest: examples & tests
make llvm # 1. in-tree LLVM/MLIR/Clang toolchain (one-time, long)
make build # 2. sar-opt, libsar_runtime, tests
make scalehls # 3. optional: ScaleHLS-HIDA toolchain for the HLS backend
export PYTHONPATH=$PWD/python:$PYTHONPATH
make test # lit + pytest, everything should pass
make examples # focus a 512x512 synthetic scene, writes a PNGThe example runners insert python/ into sys.path themselves, so they
also work without the PYTHONPATH export.
Step 3 fetches ScaleHLS' pinned LLVM via the polygeist submodule and applies two small upstream fixes (scripts/patches/).
To reproduce the San Francisco image, place the ALOS-1 CEOS product under
examples/data/ and run:
python examples/data/extract_alos.py # CEOS L1.0 -> alos_raw_16384x16384.bin
python examples/wka/run_alos.py| Document | Contents |
|---|---|
| docs/architecture.md | Layering, design decisions and their rationale |
| docs/dialect.md | The sar dialect reference: ops, passes, pipelines |
| docs/backends.md | Backend guide: cpu, scalehls, adding your own |
| examples/wka/README.md | omega-K walkthrough (synthetic + real data) |
| examples/rda/README.md | Range-Doppler walkthrough |
| examples/csa/README.md | Chirp Scaling walkthrough (interpolation-free) |
| CONTRIBUTING.md | Development setup and conventions |
include/sar/, lib/ C++ core: dialect (IR + Transforms), conversions, pipelines
runtime/ libsar_runtime (FFT, sinc interpolation; C ABI)
tools/sar-opt/ optimizer / pipeline driver
python/sar/ Python package: language, ir, compiler, runtime, backends
third_party/ backend plugins (cpu, scalehls)
examples/ omega-K, Range-Doppler and Chirp Scaling algorithms
benchmarks/ performance suite and reference numbers
test/ lit suites (MLIR) and pytest suites (Python, e2e, fuzz)
externals/ submodules: llvm-project, ScaleHLS-HIDA
scripts/ toolchain build scripts and patches
SAR-DSL is a research compiler. Current limits, by design or pending work:
- FFT sizes must be powers of two; all shapes are static (one compile per geometry).
- HLS designs are emitted and validated numerically (the same IR is compiled and checked against NumPy on the CPU), but on-device place-and-route/timing closure is left to Vitis and scene sizes are bounded by on-chip memory.
- The
cpubackend targets the host machine (JIT); it is tested on x86-64 Linux. Other Linux architectures with an LLVM host target should work but are unverified; Windows is unsupported. - Prebuilt wheels are not published; the toolchain builds from source.
Contributions are welcome — see CONTRIBUTING.md.
MIT — see LICENSE. Built on LLVM/MLIR and ScaleHLS-HIDA; sample imagery derives from JAXA ALOS-1 PALSAR data.
