claudeers.
// RAG & Knowledge

Aether

Aether — a Claude ecosystem project on GitHub.

// RAG & Knowledge[ cli ][ api ][ web ][ mobile ][ claude ]#claude#rag◷ Apache-2.0$open-sourceupdated 18 days ago
Actively maintained
100/100
last commit 17 days ago
last release 20 days ago
releases 17
open issues 0

Install with your AI

Paste into Claude Code, Cursor, or any agent — it reads the repo and wires the tool into your project.

Install and set up Aether (pip project) into my current project.
Found on https://claudeers.com/aether
Repo: https://github.com/iamkaleemsajjad-hue/Aether
Homepage/docs: —
Detected install method: pip → pip install aether-runtime
Category: rag. Platforms: cli, api, web, mobile.
Read the repo's README for exact setup and env vars, then install it and wire it into my project.

Claudeers Health Verdict:
active; community-verified: false. Confirm the source before running anything.
// or install directly (pip)
pip install aether-runtime
// or clone
git clone https://github.com/iamkaleemsajjad-hue/Aether

// compatibility

Platformscli, api, web, mobile
Operating systems—
AI compatibilityclaude
LicenseApache-2.0
Pricingopen-source
LanguagePython

Get your FREE $2.50 API credits to access TickAtlas financial data ↗

Aether Runtime

Compile once. Run on any hardware, forever.

Aether is an open-source AI model compiler and inference runtime. It ingests any open-source model (HuggingFace, GGUF, SafeTensors, ONNX) and produces a portable Aether Execution Graph (AEG) artifact that runs on any detected hardware — CPU, GPU, NPU, FPGA — with zero framework dependency and zero re-compilation.



Multi-Engine Benchmark Results — Aether vs Competitor Engines

Hardware Testbed: 2× NVIDIA Tesla T4 (14.6 GiB VRAM each) · Intel Xeon @ 2.00 GHz · FP16 Native Tensor-Core Execution · Linux 6.12 · Kaggle Environment
Software Stack: Aether Runtime v1.3.0 vs HuggingFace Transformers v5.0.0 (PyTorch 2.10.0 eager) vs PyTorch Native Decode Loop
Suite Version: 2.0.0 (Strict per-engine process isolation, identical model commits, greedy evaluation, CUDA edge synchronization)

🔗 Access Full Benchmark Results & Artifacts:


Executive Highlights — Aether Wins Across the Entire Field

Performance DimensionAether ResultCompetitor ComparisonMargin / SpeedupEvidence from Suite
Overall Win Rate100% (54 / 54)Transformers: 26% (14/54) · PyTorch Native: 0% (0/54)Undefeated across all 27 measured cells0 losses, 0 ties across the full matrix
Median Advantage+94.2% vs HFPyTorch Native: +104.3%~2x faster median throughput across all modelsPairwise anti-symmetric matrix
Peak Throughput1,562.72 tok/sTransformers: 489.17 tok/s · PyTorch Native: 476.01 tok/s3.19x faster (+219.5% margin)GPTNeo350M @ Batch 16
Batch 1 (Interactive)Swept #1, #2, #3GPTNeo: 110.77 tok/s · Qwen3: 48.21 tok/s · SmolLM2: 46.18 tok/sUp to 2.68x faster at Batch 1Aether took all top 3 spots in the field
Time-to-First-Token (TTFT)0.022s (22 ms)Transformers: 28 ms · PyTorch Native: 26 ms21% faster TTFT (Prompt tok/s: 21,016.21)SummerSigh/GPTNeo350M-Instruct-SFT
Single-Request Latency1.156sTransformers: 3.102s · PyTorch Native: 3.195s62.7% lower latency (1.95s saved per request)GPTNeo350M Batch 1 (p256 / o128)
Inter-Token Latency (TPOT)9.00 msTransformers: 24.22 ms · PyTorch Native: 24.96 ms2.7x faster per generated tokenSub-10ms token generation loop
Cold Start (Fresh Process)Swept #1, #2, #3GPTNeo: 1.415s · SmolLM2: 2.974s · Qwen3: 2.999sAll <3.0s (Competitors take 3.65s – 6.78s)First unwarmed inference in fresh process
Lowest Peak Host Memory1.615 GiBPyTorch Native: 1.677 GiB · Transformers: 1.766 GiBLowest host memory footprintSmolLM2-135M-Instruct

Overall Standings & Pairwise Head-to-Head

Every engine was scored identically by the same measurement harness across the exact same model revisions and prompt sequences:

RankEngine% of Best (Median)W / L / TWin RateMedian Diff vs FieldCells MeasuredPairings Evaluated
🥇aether100%54 / 0 / 0100%+99.2%2754 / 54
🥈transformers51%14 / 27 / 1326%-6.7%2754 / 54
🥉pytorch_native49%0 / 40 / 140%-8.7%2754 / 54

Pairwise Matrix (Median % Advantage of Row Engine over Column Engine)

Enginevs aethervs pytorch_nativevs transformers
aether—+104.3%+94.2%
pytorch_native-51.0%—-2.0%
transformers-48.5%+2.0%—

Visual Benchmark Comparisons — 3 Engines across 3 Architectures

The charts below illustrate empirical measurements extracted directly from the comprehensive benchmark suite across all three architectures (SmolLM2-135M, GPTNeo-350M, Qwen3-0.6B) executed under strictly identical conditions on 2× NVIDIA Tesla T4 GPUs (FP16 Native Tensor-Core Execution).

1. Output Tokens Per Second (Throughput) — 3 Engines across 3 Models

Output Tokens Per Second Comparison

Comparison of output tokens per second across all 3 evaluated models: Single-Request Interactive Throughput (Batch 1, prompt=256, output=128, left panel) and Peak Batched Serving Throughput (Batch 16, prompt=256, output=128, right panel). Aether delivers 1.71x to 2.68x (+70.7% to +168.4%) higher single-stream throughput and scales up to 1,562.72 tok/s at Batch 16.

2. End-to-End Single-Request Latency (Not TTFT) — 3 Engines across 3 Models

End-to-End Latency Comparison

Single-request end-to-end latency (seconds per request for 128 generated tokens; ▼ lower is better). This measures full decode request completion time—distinct from Time-To-First-Token (TTFT)—where Aether decisively outperforms the field across all three architectures, cutting latency by 41.4% to 62.7% and saving 1.95s to 2.91s per request compared to HuggingFace Transformers and PyTorch Native.


Model-by-Model Results with Exact Empirical Evidence

1. SummerSigh/GPTNeo350M-Instruct-SFT (456M Params)

Decisive Wins: Peak throughput reached 1,562.72 tok/s (+219.5% margin over Transformers). Single-user Batch 1 throughput reached 110.77 tok/s (2.68x faster than Transformers at 41.27 tok/s) with a 62.7% reduction in end-to-end latency (1.156s vs 3.102s). TTFT dropped to 22 ms with prefill throughput exceeding 21,016 prompt tok/s.

Throughput & Scaling Across Batch Sizes (Prompt: 256, Output: 128)
Batch SizeAether (tok/s)Transformers (tok/s)PyTorch Native (tok/s)Aether vs HF SpeedupAether vs PyTorch SpeedupScaling Efficiency
b1110.7741.2740.062.68x (+168.4%)2.77x (+176.5%)100%
b2206.7983.8281.902.47x (+146.7%)2.52x (+152.5%)93%
b4447.40165.86161.172.70x (+169.7%)2.78x (+177.6%)101%
b8849.21303.24295.322.80x (+180.0%)2.88x (+187.6%)96%
b161,562.72489.17476.013.19x (+219.5%)3.28x (+228.3%)88%
Prompt & Output Length Sweeps at Batch 1
Prompt TokensOutput TokensAether (tok/s)Aether LatencyHF (tok/s)HF LatencyPyTorch Native (tok/s)Aether Margin
32128112.921.134s39.603.232s40.04+185.1% (2.85x)
25632102.200.313s40.880.783s39.22+150.0% (2.50x)
256128110.771.156s41.273.102s40.06+168.4% (2.68x)
256512116.764.385s41.0512.473s40.42+184.5% (2.84x)
1024128116.151.102s39.783.218s39.19+192.0% (2.92x)

2. Qwen/Qwen3-0.6B (752M Params — RoPE + Per-Head Q/K Norm)

Decisive Wins: Batch 1 throughput achieved 48.21 tok/s (more than double Transformers' 23.01 tok/s, a +109.5% margin). Request latency dropped from 5.56s to 2.65s (52.3% lower). Inter-token latency improved from 43.42 ms down to 20.63 ms/token. First-call cold start took only 2.99s compared to Transformers' 6.14s.

Throughput & Scaling Across Batch Sizes (Prompt: 256, Output: 128)
Batch SizeAether (tok/s)Transformers (tok/s)PyTorch Native (tok/s)Aether vs HF SpeedupAether vs PyTorch Speedup
b148.2123.0121.932.10x (+109.5%)2.20x (+119.8%)
b282.6344.8643.391.84x (+84.2%)1.90x (+90.4%)
b4165.5488.9685.601.86x (+86.1%)1.93x (+93.4%)
b8234.55171.74165.881.37x (+36.6%)1.41x (+41.4%)
b16273.35241.03239.231.13x (+13.4%)1.14x (+14.3%)
Prompt & Output Length Sweeps at Batch 1
Prompt TokensOutput TokensAether (tok/s)Aether LatencyHF (tok/s)HF LatencyPyTorch Native (tok/s)Aether Margin
3212848.342.648s22.535.681s21.89+114.6% (2.15x)
2563244.300.722s22.811.403s21.69+94.2% (1.94x)
25612848.212.655s23.015.563s21.93+109.5% (2.10x)
25651250.3710.164s22.4922.766s22.22+124.0% (2.24x)
102412846.102.777s22.115.788s21.52+108.5% (2.09x)

3. HuggingFaceTB/SmolLM2-135M-Instruct (135M Params)

Decisive Wins: Smooth scaling from 46.18 tok/s at Batch 1 to 674.99 tok/s at Batch 16 (+63.4% margin over Transformers). Interactive latency improved from 4.73s to 2.77s (41.4% faster). Peak host resident memory was only 1.615 GiB (lowest of any engine). Prompt processing speed reached 8,905.61 prompt tok/s (+44.2% faster prefill).

Throughput & Scaling Across Batch Sizes (Prompt: 256, Output: 128)
Batch SizeAether (tok/s)Transformers (tok/s)PyTorch Native (tok/s)Aether vs HF SpeedupAether vs PyTorch Speedup
b146.1827.0627.321.71x (+70.6%)1.69x (+69.1%)
b285.7452.9152.901.62x (+62.0%)1.62x (+62.1%)
b4178.19105.85105.231.68x (+68.3%)1.69x (+69.3%)
b8355.94210.49208.091.69x (+69.1%)1.71x (+71.0%)
b16674.99413.00410.201.63x (+63.4%)1.65x (+64.6%)
Prompt & Output Length Sweeps at Batch 1
Prompt TokensOutput TokensAether (tok/s)Aether LatencyHF (tok/s)HF LatencyPyTorch Native (tok/s)Aether Margin
3212844.932.849s27.244.700s27.30+65.0% (1.65x)
2563242.750.749s27.161.178s26.62+57.4% (1.57x)
25612846.182.772s27.064.729s27.32+70.6% (1.71x)
25651248.0710.651s27.0518.927s27.45+77.7% (1.78x)
102412846.942.727s26.594.814s27.00+76.6% (1.77x)

Why Aether Outperforms: Architectural & Compilation Advantage

  1. AOT Ahead-of-Time Graph Compilation: Eliminates Python interpreter overhead and PyTorch dynamic dispatch loops during token generation.
  2. Fused Custom Kernels: Native C++ kernels executing fused RMSNorm + SwiGLU / GeGLU and FlashAttention-2 paths optimize memory bandwidth and reduce device kernel launches.
  3. Optimized KV-Cache Layout: Zero-copy continuous memory buffers prevent cache fragmentation and preserve memory bandwidth under scaling.
  4. Compilation Amortization: On SmolLM2-135M, Aether's 7.0s AOT compilation saves 1.96 seconds on every subsequent generation request — breaking even and pulling permanently ahead after just 4 inference requests.

💡 View the complete suite run, raw measurements, and all 31 charts in benchmark/results/benchmark_results.ipynb and benchmark/results/BENCHMARK_RESULTS.md.


Core Principles

PrincipleImplementation
Compile once, run anywhereAEG artifacts are hardware-portable; they contain multi-target sharding plans and run without re-compilation
PyTorch-free coreThe runtime, compiler, and CPU engine require only NumPy + tokenizers. PyTorch is optional (pip install "aether-runtime[pytorch]")
Universal hardware detectionDetects NVIDIA (CUDA), AMD (ROCm), Apple (Metal/MPS), Intel (OpenVINO), Qualcomm (QNN), RISC-V, FPGA, and pure CPU — no driver installation required
Multi-GPU with VRAM-weighted distributionAutomatically shards model weights across all available GPUs proportional to each GPU's VRAM capacity
Framework-free native kernelsC++ kernels compiled at runtime: INT4-GEMV, FlashAttention-2, fused RMSNorm+SwiGLU+Linear, GeGLU, RoPE, OpenMP parallel SGEMM

5-Stage Compiler Pipeline

Model (HuggingFace / GGUF / SafeTensors / ONNX)
        │
        ▼
┌─────────────────────────────────────────────────────────────────┐
│ Stage 1: Ingestion & Architecture Detection                     │
│  • Reads config.json / GGUF header / SafeTensors metadata       │
│  • Detects 60+ model families without relying on model names    │
│  • Outputs: AEG-IR computation graph + ModelArchitecture        │
└─────────────────────────────────────────────────────────────────┘
        │
        ▼
┌─────────────────────────────────────────────────────────────────┐
│ Stage 2: Optimizer (22 Passes)                                  │
│  Pass 1: Operator Fusion (RMSNorm→QKV→RoPE, SwiGLU fusion)     │
│  Pass 2: Sensitivity Analysis (per-layer perplexity gradient)   │
│  Pass 3: Precision Assignment (mixed-precision per sensitivity) │
│  Pass 4: KV Cache Structuring (paged blocks, radix-tree hints)  │
│  Pass 5: MoE Expert Routing (hot/warm/cold tier classification) │
│  Pass 6: Parallelism Discovery (TP/PP/EP/CP strategy search)   │
│  Pass 7–22: Graph lowering, sparse attention, pruning, etc.     │
└─────────────────────────────────────────────────────────────────┘
        │
        ▼
┌─────────────────────────────────────────────────────────────────┐
│ Stage 3: Quantization                                           │
│  • Q4_K_M (INT4 block-scaled) — default for ≤70B models        │
│  • Q8_0 (INT8 symmetric)                                        │
│  • BF16 / FP16 / FP8 (E4M3 / E5M2)                            │
│  • MXFP4 / MXFP6 (microscaling, PRD v4.0+)                     │
└─────────────────────────────────────────────────────────────────┘
        │
        ▼
┌─────────────────────────────────────────────────────────────────┐
│ Stage 4: AEG Packaging                                          │
│  • Self-contained .aeg/ directory with manifest, weights,       │
│    tokenizer, precision map, sharding plans                     │
│  • Integrity-verified (SHA-256 per artifact)                    │
│  • Version-stamped (AEG/1.1 – AEG/3.0)                        │
└─────────────────────────────────────────────────────────────────┘
        │
        ▼
┌─────────────────────────────────────────────────────────────────┐
│ Stage 5: Target Code Generation                                 │
│  • CUDA (sm70–sm130), ROCm (RDNA3, CDNA3/4/5)                 │
│  • Apple Metal (M1–M5), OpenVINO (NPU/GPU)                     │
│  • Qualcomm QNN, RISC-V NPU, FPGA                              │
│  • Native CPU (AVX-512, AVX2, NEON, ternary BitNet)            │
└─────────────────────────────────────────────────────────────────┘
        │
        ▼
  AEG artifact (.aeg/) — runs anywhere, forever

Quick Start

pip install aether-runtime

# Compile a model to AEG
aether compile meta-llama/Llama-3.1-8B --target cuda_sm90 --precision q4_k_m

# Inspect the compiled artifact
aether inspect llama-3.1-8b.aeg/

# Run inference (no GPU required for CPU target)
aether serve llama-3.1-8b.aeg/ --port 8080

# Benchmark performance
aether bench llama-3.1-8b.aeg/

# Run evaluation gate
aether eval llama-3.1-8b.aeg/ --suite reasoning --max-regression 0.02

Python API

from aether.compiler import AetherCompiler
from aether.compiler.config import CompilerConfig

# Compile — no PyTorch required
compiler = AetherCompiler()
config = CompilerConfig(
    target="cuda_sm90",
    precision="q4_k_m",
    max_context_length=131072,
)
artifact = compiler.compile("meta-llama/Llama-3.1-8B", config)
print(artifact.summary())   # model_id, target, precision, size_gb, tok/s estimate

# Run compiled AEG
from aether.backends import get_backend

backend = get_backend("aether_cpu")   # or "vllm", "mlx", "onnxruntime"
backend.load_model("llama-3.1-8b", aeg_path="llama-3.1-8b.aeg/")
result = backend.generate(GenerationRequest(
    model_id="llama-3.1-8b",
    prompt="What is quantum entanglement?",
    max_tokens=512,
))
print(result.text)
print(f"Throughput: {result.metrics['throughput_tps']:.1f} tok/s")

Multi-GPU Execution (Hardware-Aware Placement Planner)

When more than one device is present, Aether does not assume it should use them. It plans: it measures the machine, reads the model's exact tensor geometry out of the AEG, and judges every structurally admissible placement on two separate axes.

aether plan model.aeg --batch 4 --context 8192 --intent balanced
FEASIBILITY  binding device     1x cuda:0      TP=2/cap      PP=2/bal
  C_safe           GiB           12.96         12.96         12.96
  static S         GiB           15.51          7.88          7.88
  transient T      GiB            0.51          0.44          0.44
  margin z*sigma    GiB           0.17          0.15          0.15
  KV budget K      GiB           -3.23          4.49          4.49
  tokens_max                          0        74,724        74,724
  verdict                    INFEASIBLE      feasible      feasible

PERFORMANCE  decode/token     1x cuda:0      TP=2/cap      PP=2/bal
  bandwidth roof    ms           51.20         25.60         51.20
  dispatch roof     ms           27.16         59.75         27.16
  predicted TPOT    ms           51.20         61.16         51.22
  binding roof                bandwidth      dispatch     bandwidth

Feasibility is a residual, not a comparison. KV cache is the elastic term, so it is what is left over:

C_safe(d) = min(free(d) - external(d), total(d)*kappa) - R_fixed(d)
K(d)      = C_safe(d) - static(d) - (transient(d) + z*sigma(d))
tokens_max = min_d floor(K(d) / kv_per_token(d))

That turns "does it fit" into a capacity, which is what lets Aether answer "batch size just changed" without replanning and report the context ceiling at load time instead of discovering it as an OOM. kappa is the only percentage in the model, and its job is to absorb driver growth — not to test fit. R_fixed (CUDA context, cuBLAS workspace, collective buffers) is measured, not modelled, and sigma is the standard deviation of this device's own past prediction errors, so the safety margin shrinks as evidence accumulates.

Performance is three roofs, not two. A Python-dispatched runtime has a ceiling the roofline model does not contain, and for small-model decode it is the binding one:

t_stage = max( FLOPs/(theta_flops*u) , bytes/theta_bw , n_ops*t_dispatch )

Aether measures Qwen3-0.6B at 41.96 tok/s on one T4 — 23.8 ms/token. The two-roof model predicts 3.75 ms and therefore recommends sharding. The model was never near its bandwidth roof, and TP roughly doubles the host op count: predicted 53 ms, measured ~2× slower. The third roof is what makes the planner get this right, and it also produces the useful advice — dispatch-bound at 23.8 ms against a 3.8 ms bandwidth roof; capture CUDA graphs, don't add GPUs.

Two structural laws prune the search before any ranking happens, which is why planning takes under a millisecond for 8 devices instead of Gurobi-hours:

  • Homogeneity — a TP group's devices must be within a derived throughput ratio of each other: max(1 + sigma, heads*sigma - 1), the wider of the throughput-measurement noise floor and the ratio at which rounding a shard to a whole attention head breaks the planner's own error bar. A TP group is a barrier twice per layer, so the slowest member sets the pace on every layer. The bound tightens as calibration accumulates, and a measured crossover — recorded whenever a heterogeneous group misses its water-filled prediction — overrides the derivation outright. CPUs and mismatched GPUs never join one, by arithmetic rather than by name.
  • Fabric alignment — a TP group may not cross a fabric class. Heterogeneity is expressed across pipeline stages, where each runs at its own pace.

Both laws are structural: they ask whether a group can be balanced, never whether widening is worthwhile. That question belongs to the ranking lane, and keeping it there is what stops the generator from deleting the only plan a too-large model has.

Asymmetric splits are water-filling, and the objective is phase-dependent. For a 16 GB + 24 GB pair holding a 21 GiB model, the capacity-optimal split is 33.8 / 66.2 — which holds 74,724 KV tokens against 30,425 for a naive 50/50. Not 50/50, not memory-proportional, not bandwidth-proportional: the constrained optimum.

SizingObjectiveRule
TP shard fractionsmin max t_iwater-fill ∝ θ, capped
PP layers, throughputmin max t_iwater-fill ∝ θ, capped
PP layers, latencymin Σ t_igreedy — fastest device first

A tie goes to fewer devices. A wider plan is accepted only when its predicted gain exceeds the planner's own error bar, which makes "use the minimum hardware necessary" a consequence of the cost model rather than a preference. The same 34B model on two NVLink A100s is selected as TP=2 (1.97× faster, 4.6× the KV) under a graph-captured runtime and 1× cuda:0 under an eager one — one formula, opposite answers, both correct.

Nothing is reactive: there is no OOM-and-retry path. An impossible workload is refused before the load with the arithmetic and the fixes that would change the answer. Telemetry feeds a calibration ledger keyed by device signature and backend build, so the next prediction is tighter — never the current placement.

The first run calibrates itself. With no ledger entry there is no measured sigma, so the planner runs one forward pass at the workload ceiling after the weights are resident, reads peak allocated and cuda_used - torch_reserved, and folds both in — one profile run for the chosen plan, not one per candidate. An allocation failure during that pass is recorded as evidence the prediction was low rather than raised, because the pass exists to protect the process. Set AETHER_PLAN_BOOTSTRAP=0 to skip it; the record then says the margin is uncalibrated instead of pretending otherwise.

t_dispatch is verified, not trusted. It belongs to the runtime build, so its ledger key carries the interpreter, the framework and its CUDA build, Aether's own version and the execution mode — and because no key can capture every change, a fresh probe reconciles against the stored value on every census and replaces it when they diverge by 2×. The record also prints the dispatch cost at which the verdict would flip ("PP=2 would win above 44.1 us/op"), so a mis-keyed value cannot bias the answer silently even if it slips past both defences.

Full design, including the fourteen stress-tested scenarios and what would falsify it: docs/architecture-execution-planner.html. Implementation: src/aether/placement/.

from aether.placement import ExecutionPlanner, Intent, WorkloadEnvelope
from aether.placement.model_profile import profile_from_manifest

planner = ExecutionPlanner(profile_from_manifest(manifest))
decision = planner.plan(WorkloadEnvelope(
    batch_target=4, context_target=8192, generate_target=512, intent=Intent.BALANCED,
))
print(decision.render())                    # the full derivation
print(decision.selected.tokens_max)         # capacity, not a boolean
print(decision.plan.device_ids)             # which devices, and why

The compiler embeds sharding plans for 1–8 GPUs in every AEG artifact (Pass 6: Parallelism Discovery). At runtime the distributed engine reads the matching plan and reduces with Aether's own collectives.

Which collective runs where — precisely. "No NCCL" is true of two of the three paths, and the difference matters:

Execution modeCollectiveNCCL / torch.distributed?
CPU, multi-processSocketCollective — ring reduce-scatter + all-gather over TCPNot required. Verified across real processes up to 8 ranks.
Single-process, multi-GPUaether.parallelism.p2p_ring — one-shot / two-shot / ring over CUDA-ROCm peer-to-peer device copiesNot required. This is the path the tensor-parallel executor uses.
Multi-process or multi-node GPUNCCL (CUDA) or RCCL (ROCm) via torch.distributedRequired. Aether does not reimplement inter-node GPU transport, and asking for this backend on a host without it fails closed.

The peer-to-peer path picks its schedule per call from the α–β cost model, using the detected link latency and bandwidth — because no single schedule is right at both ends of the size range:

one-shot   α + (P−1)·D/B            volume (P−1)·D        — small payloads, latency-bound
two-shot   2α + 2(P−1)/P·D/B        volume 2(P−1)/P·D     — large payloads, fully peer-connected
ring       2(P−1)·α + 2(P−1)/P·D/B  volume 2(P−1)/P·D     — meshes without full peer access

Crossover, from setting the first two equal: D* = α·P·B / ((P−1)(P−2)) for P > 2. At P = 2 one-shot is never worse — same volume, half the hops.

Every collective fails closed. A ring that loses a peer raises CollectiveError rather than returning an approximation. Reductions run in a fixed device order, so results are bit-reproducible and every device gets identical bytes.

from aether.parallelism.p2p_ring import P2PRingCollective

collective = P2PRingCollective(["cuda:0", "cuda:1", "cuda:2", "cuda:3"])
reduced = collective.all_reduce(per_device_partials)   # every device: the full sum
root    = collective.reduce_to_root(per_device_partials)  # tree, ceil(log2 P) rounds
print(collective.stats()["requires_nccl"])            # False

References: Patarasuk & Yuan, JPDC 69(2), 2009 (ring bandwidth bound); Thakur, Rabenseifner & Gropp, IJHPCA 19(1), 2005 (algorithm choice by message size); Shoeybi et al., arXiv:1909.08053 §3.3 (why this all-reduce dominates TP cost).


Hardware Detection

from aether.backends.hardware_detector import detect_hardware

profile = detect_hardware()
print(profile.summary())
# ┌─────────────────────────────────────────────────────┐
# │ Aether Hardware Profile                             │
# │  GPUs:  2x NVIDIA RTX 4090 (24 GB each) → cuda_sm89│
# │  CPU:   AMD Ryzen 9 7950X  (AVX-512)   → cpu_avx512│
# │  RAM:   128 GB                                      │
# │  Best target: cuda_sm89                             │
# └─────────────────────────────────────────────────────┘

Detected targets span: NVIDIA CUDA (sm70–sm130), AMD ROCm (RDNA3, CDNA3/4/5), Apple Metal (M1–M5), Intel OpenVINO (NPU/GPU), Qualcomm QNN, RISC-V NPUs (SiFive X160, XuanTie C930, MIPS S8200), Xilinx FPGA, and CPU (AVX-512, AVX2, NEON, ternary).


Native CPU Kernel Stack

The Aether CPU engine compiles a C++ shared library at first run (cached thereafter) providing the following kernels — all OpenMP-parallel and auto-vectorized to AVX-512/NEON:

KernelDescriptionResearch Basis
aether_int4_gemvINT4-packed GEMV — 2x bandwidth vs INT8, ~2x tok/sGGML Q4_0 (2023), GPTQ (2022)
aether_qgemv_affineINT8 affine-quantized GEMVFrantar et al. 2022
aether_flash_attnFlashAttention-2 online softmax (O(seq·d) memory)Dao, NeurIPS 2023
aether_rmsnorm_linearFused RMSNorm + QKV projection (1 buffer)ClusterFusion NeurIPS 2025
aether_rmsnorm_swiglu_linearFused RMSNorm + full SwiGLU FFN (0 intermediate buffers)Shazeer 2020, ClusterFusion 2025
aether_gegluGeGLU for Gemma/Gemma-2 FFNHendrycks 2016, Google Gemma 2024
aether_swigluSwiGLU activation (Llama/Qwen/Mistral)Shazeer 2020
aether_ropeRotary position embedding in-placeSu et al. 2021
aether_sgemvFP32 GEMV (M=1 decode fast path, ~3x vs SGEMM)BLIS (Van Zee 2015)
aether_sgemmCache-blocked FP32 SGEMM (prefill)BLIS tile layout
aether_softmaxNumerically stable row-wise softmaxStandard
aether_rmsnormDouble-accumulation RMSNormZhang & Sennrich 2019
aether_argmaxGreedy token selection (OpenMP reduction)Standard

No compiler toolchain? Every kernel has a NumPy fallback — the module always imports and runs.


Supported Model Families

Aether classifies 40 architecture families, reached through 164 model-name and Hugging Face architecture-class spellings. Support is graded, because "it runs" and "its logits match the reference" are different claims:

LevelFamiliesMeaning
✅ Parity-verified26Every logit compared against the 🤗 Transformers reference (~1e-6) on the CPU, PyTorch, and tensor-parallel engines, for prefill and decode
🟡 Runs, not gated6Compile → load → execute round-trip tested; no automatic per-logit comparison yet (5 encoders + T5/BART-class seq2seq)
🔬 Known-incorrect4Executes, but measured output diverges from the reference — documented, not relied upon (Mamba, Mamba-2, RWKV-7, Jamba)
❌ Refused4Detected and then rejected at compile time rather than producing a wrong artifact (DeepSeek MLA, MiniMax, VLM, Whisper)

36 families are executable; 26 are verified. The exact numbers come from src/aether/core/model_families.py and are asserted against this table by tests/unit/test_model_family_registry.py, so they cannot drift. Print them yourself:

aether models              # the full graded matrix
aether models --counts     # just the numbers

A fine-tune of a verified family is covered by that family — detection keys on structure, not on name — which is why Vicuna, Zephyr, Dolphin, Tulu, Nous-Hermes, OpenChat, TinyLlama, Yi, InternLM, MiniCPM, SOLAR and the rest of the Llama/Qwen/Mistral derivative space add detection keys rather than families. See SUPPORTED_MODELS.md for the per-family matrix with the distinguishing numerics Aether derives from each checkpoint.

The 26 parity-verified families

FamilyModelsDistinguishing contract
Llama 3.xLlama-3.1-8B, 3.2-1B/3B, 3.3-70BGQA + SwiGLU + RMSNorm + RoPE (baseline)
Qwen 2 / 2.5Qwen2-7B/72B, Qwen2.5, CodeQwenGQA + SwiGLU, schedule-gated sliding window
Qwen 3Qwen3-0.6B → 72Bper-head Q/K-norm, decoupled head_dim
Qwen 3 MoEQwen3-MoEexperts without top-k renormalization
MistralMistral-7B v0.1–v0.3, MinistralGQA + SwiGLU
MixtralMixtral-8x7B/8x22Btop-2 of 8 experts with renormalization
Gemma 2Gemma-2-2B/9B/27B×√H embeddings, (1+w) norms, sandwich norm, logit soft-caps, GeGLU
Gemma 3 (text)Gemma-3-1B/4B/12B/27BGemma 2 + separate local rotary base
GPT-2GPT-2 117M–1.5B, DialoGPTConv1D layout, GELU-tanh, learned positions
GPT-NeoGPT-Neo 125M/1.3B/2.7Bunscaled attention, local/global schedule
GPT-NeoXGPT-NeoX-20B, Pythia 70M–12B25% partial rotary, head-interleaved QKV, parallel residual
GPT-JGPT-J-6Binterleaved rotary, parallel residual
Phi-3 / Phi-4Phi-3-mini/small/medium, Phi-4fused QKV, LongRoPE factor tables
FalconFalcon-7B/40Bper-KV-group interleaved QKV, parallel residual
BLOOMBLOOM 560M–176B, BLOOMZALiBi, embedding LayerNorm
MPTMPT-7B/30BALiBi, nested attn_config spellings
StarCoder2StarCoder2-3B/7B/15BGQA, GELU-tanh, layer_types window
Cohere / Command-RCommand-R/R+/A, Aya Expanseinterleaved rotary, logit_scale
OLMo 2OLMo-2-7B/13Bpost-norm block, full-projection Q/K-norm
OLMoEOLMoE-1B-7Bfull-projection Q/K-norm, unnormalized experts
StableLMStableLM-2, StableLM-3B25% partial rotary
GraniteGranite-3.x, Granite Codeembedding/residual/attention/logit multipliers
EXAONE 4EXAONE-4-32Bpost-norm, NoPE global layers
SmolLM 3SmolLM3-3Binterleaved NoPE layers
GLM-4GLM-4-9B/32Binterleaved + 50% partial rotary, GLM sandwich norm
NemotronNemotron-4, Nemotron-MiniLayerNorm1P, squared-ReLU FFN

Supported Hardware Targets

Target IDHardware
cuda_sm89NVIDIA RTX 4090 (Ada Lovelace)
cuda_sm90NVIDIA H100 (Hopper)
cuda_sm100NVIDIA B200 (Blackwell)
cuda_sm130NVIDIA Rubin Ultra (sm_130)
rocm_cdna3AMD MI300X
rocm_cdna5_mi455xAMD MI455X (CDNA5)
metal_m3Apple M3/M4/M5
openvino_npuIntel Arc NPU
qualcomm_qnnQualcomm Snapdragon NPU
cpu_avx512x86-64 with AVX-512
cpu_neonARM NEON (mobile, Raspberry Pi)
cpu_avx512_ternaryBitNet b1.58 ternary on x86
riscv_sifive_x160SiFive Intelligence X160
fpga_xilinx_vu9pXilinx VU9P (decode-only)

Full list: 30+ targets in src/aether/core/constants.py.


Phase 5 — Observability

Aether emits OTLP directly — the wire protocol, not a JSON file that resembles it. trace_id/span_id widths, typed AnyValue attributes (intValue as a string, per the protobuf JSON mapping), timeUnixNano events, per-span kind, real gzip when Content-Encoding: gzip is advertised, W3C traceparent propagation, trace-ID-ratio sampling, and the standard OTEL_* environment variables. No dependency is needed for any of it; conformance is pinned by tests/unit/test_otlp_conformance.py.

from aether.observability.otel import AetherTracer, OTLPExporter, MetricsCollector

tracer = AetherTracer(service_name="aether-prod", sample_rate=0.01)
exporter = OTLPExporter()          # honours OTEL_EXPORTER_OTLP_ENDPOINT/HEADERS/TIMEOUT
exporter.export_to_endpoint(tracer)

metrics = MetricsCollector()
exporter.export_metrics_to_endpoint(metrics)   # real explicit-bucket histograms

Joining a trace that started upstream, and a span that records its own failure:

with tracer.span("aether.prefill", traceparent=request.headers.get("traceparent")) as span:
    span.add_event("kv_built", {"blocks": 128})

Routing through an existing OpenTelemetry SDK pipeline — span processors, resource detectors, propagators, exporters configured by the host application — is the one thing that needs the dependency:

pip install "aether-runtime[otel]"
from aether.observability.otel_sdk import OpenTelemetryBridge, is_available

if is_available():
    OpenTelemetryBridge("aether-prod").emit_all(tracer.get_finished_spans())
    # spans keep Aether's trace_id, so they correlate rather than duplicating
from aether.observability.ci_pipeline import CIEvalPipeline
from aether.observability.gates import DriftMonitor, ABRolloutController

pipeline = CIEvalPipeline(aeg_path='model.aeg', max_regression=0.02)
report = pipeline.run_and_save('eval_report.json', benchmarks=['hellaswag', 'mmlu', 'gsm8k'])

ctrl = ABRolloutController('exp-001', candidate_percent=0.01)
monitor = DriftMonitor(baseline_win_rate=0.80, alert_drop=0.05, min_samples=20)

Prometheus metrics: aether_request_total, aether_ttft_ms{quantile=p50|p95|p99}, aether_tokens_per_second, aether_kv_hit_rate, aether_eagle_accept_rate


Content Credentials (C2PA)

aether sign writes a real C2PA manifest store to provenance/c2pa.manifest inside the package — not a hash chain with C2PA-shaped field names:

  • a c2pa.claim.v2 claim in deterministic CBOR (RFC 8949 §4.2.1), referencing every assertion by hashed URI;
  • an assertion store with the hard binding, c2pa.actions.v2, c2pa.ingredient.v3 for the source checkpoint, and the compiler-pass chain;
  • a COSE_Sign1 claim signature (RFC 9052), detached, with the signer's X.509 chain in the x5chain protected header;
  • the tree serialized as JUMBF boxes (ISO/IEC 19566-5);
  • a c2pa.hash.collection.data hard binding — one digest per file, so verification reports which file changed.

Ed25519 (RFC 8032), CBOR, COSE and JUMBF are implemented in pure Python, so signing works on a stock CPython install; cryptography is used when present for speed and for the ECDSA algorithms.

aether sign   ./model.aeg                      # generates a key on first use
aether verify ./model.aeg                      # exits non-zero if integrity fails
aether verify ./model.aeg --trust-anchor ca.pem

Verification reports five checks independently — manifest present, structure, claim signature, assertion hashes, file binding — because the failures mean different things. Integrity is not identity: a self-signed manifest proves the artifact is unmodified and says nothing about who produced it, and aether verify states that rather than printing "verified".


Phase 6 — Ecosystem

from aether.ecosystem.sdks import TypeScriptSDKGenerator, GoSDKGenerator, RustSDKGenerator

TypeScriptSDKGenerator().write('./sdk/typescript/')   # aether-sdk.ts
GoSDKGenerator().write('./sdk/go/')                   # aether_client.go
RustSDKGenerator().write('./sdk/rust/src/')           # aether_client.rs

CLI Reference

CommandDescription
aether compile <model>Compile model to AEG package
aether inspect <path.aeg>Show AEG package summary
aether bench <path.aeg>Run benchmark suite
aether serve <path.aeg>Start inference server
aether eval <path.aeg>Run eval gate CI check
aether hardwareShow hardware profile
aether modelsShow the graded model-family support matrix
aether plan <path.aeg>Show the hardware-aware placement decision and its derivation
aether hub push <path.aeg>Push to Aether Hub CDN
aether hub pull <model-id>Pull from Aether Hub CDN
aether sdk generateGenerate TypeScript/Go/Rust SDKs
aether sign <path.aeg>Sign with C2PA Content Credentials (CBOR claim + COSE_Sign1 + JUMBF)
aether verify <path.aeg>Verify the claim signature, assertion hashes, and per-file binding

Installation

# Core runtime (no PyTorch, no CUDA required)
pip install aether-runtime

# With PyTorch for .pt/.pth model ingestion
pip install "aether-runtime[pytorch]"

# With HuggingFace Transformers for AutoTokenizer, AutoConfig
pip install "aether-runtime[transformers-frontend]"

# Full install (all optional backends)
pip install "aether-runtime[full]"

Testing

python -m pytest tests/ -v                                     # All tests
python -m pytest tests/unit/test_native_cpu_kernels.py -v     # CPU kernels
python -m pytest tests/unit/test_phase5_observability.py -v   # Observability
python -m pytest tests/unit/test_phase6_ecosystem.py -v       # Ecosystem SDKs
python -m pytest tests/unit/test_v31_elite_extensions.py -v   # v3.1 extensions

Run the suite serially. test_e2e_compile_run_cpu.py and test_v31_features.py share the ~/.aether cache; parallel pytest workers race on it.

Tests requiring HuggingFace weights skip cleanly when offline.


Research Citations

FeatureResearch
INT4 GEMVGerganov GGML Q4_0 (2023), Frantar GPTQ (2022)
FlashAttention-2Dao et al., NeurIPS 2023
Operator FusionClusterFusion, NeurIPS 2025
SwiGLU / GeGLUShazeer 2020 (GLU Variants); Hendrycks & Gimpel 2016
BLIS SGEMM tilesVan Zee & van de Geijn, TOMS 2015
VRAM-weighted TPMegatron-LM (Shoeybi et al. 2019); DeepSpeed (Rasley et al. 2020)
MoE Expert RoutingZipf prior: Zoph et al. 2022; Fedus et al. 2022
Sparse AttentionMInference (Microsoft, NeurIPS 2024)
KV EvictionStreamingLLM (2023), ScissorHands (2024), SnapKV (2025)
Ring AttentionRing Attention (2023), Striped Attention (2023)
YaRN RoPEYaRN (2023), LongRoPE (2024)
Speculative DecodingEAGLE-2 (2024), Medusa (2024)
CUDA GraphsvLLM CUDA Graphs Dispatcher (2026)
Process Reward ModelLet's Verify Step by Step (2023), OmegaPRM (2025)
IP FingerprintingMetaFinger (2024), ADV-TRA (2025)
EU AI Act ComplianceArticle 50 — AI content transparency obligations
Fleet SchedulingHelium (2026), MuxWise SLO-aware scheduling (2026)
Disaggregated ServingDistServe (2024), Mooncake (2024)

License

Apache 2.0

Aether Runtime — Compile once. Run on any hardware, forever.

// faq

What is Aether?

Aether is a Claude ecosystem project. It is open-source on GitHub.

Is Aether free to use?

Aether is open-source under the Apache-2.0 license, so it is free to use.

What category does Aether belong to?

Aether is listed under rag in the Claudeers registry of Claude-compatible tools.

8 views
★ 77 stars
unclaimed
updated 18 days ago

// embed badge

Aether on Claudeers
[![Claudeers](https://claudeers.com/api/badge/aether.svg)](https://claudeers.com/aether)

// retro hit counter

Aether hit counter
[![Hits](https://claudeers.com/api/counter/aether.svg)](https://claudeers.com/aether)

// reviews

// guestbook

0/500

// related in RAG & Knowledge

🔓

Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant contex…

// ragthedotmack/⟨JavaScript⟩★ 95,595◷ Apache-2.0[ claude ]
🔓

✨ Light and Fast AI Assistant. Support: Web | iOS | MacOS | Android | Linux | Windows

// ragChatGPTNextWeb/⟨TypeScript⟩★ 88,803◷ MIT[ claude ]
🔓

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server.

// ragheadroomlabs-ai/⟨Python⟩★ 73,654◷ Apache-2.0[ claude ]
🔓

A light-weight and powerful meta-prompting, context engineering and spec-driven development system for Claude Code by TÂCHES.

// raggsd-build/⟨JavaScript⟩★ 64,462◷ MIT[ claude ]
→ see how Aether connects across the ecosystem