Graviton streams model layers from SSD and quantizes in-flight (FP16 to 4-bit/2-bit/1.58-bit) so models never need to fit in memory all at once. A 72B model compresses to 36 GB and runs on a 64 GB Mac. Includes speculative decoding for 2-3x throughput, dynamic sparsity, REST API for agents, and a chat UI.
- FREE TIER
- Completely free and open-source under Apache 2.0 license.
- COMPANY
- OpenGraviton
- CATEGORY
- Infrastructure › Local Inference
- OPEN SOURCE
- Yes
- PRIVACY
- Runs entirely locally, no data leaves your machine
- TAGS
- local-llmquantizationinferenceapple-siliconopen-sourceself-hosted
BEST FOR
- ▸Users with consumer hardware running huge LLMs locally, even 500B+.
- ▸Developers building private AI agents; offers local REST API & performance.
- ▸Privacy-focused users requiring powerful LLMs without data ever leaving device.
NOT FOR
- ✕Users seeking a simple, zero-setup cloud-based LLM experience.
- ✕Non-technical users expecting a plug-and-play chat interface without setup.