Graviton

Graviton streams model layers from SSD and quantizes in-flight (FP16 to 4-bit/2-bit/1.58-bit) so models never need to fit in memory all at once. A 72B model compresses to 36 GB and runs on a 64 GB Mac. Includes speculative decoding for 2-3x throughput, dynamic sparsity, REST API for agents, and a chat UI.

AI ToolFree Tier: Completely free and open-source under Apache 2.0 license.Company: OpenGravitonCategory: InfrastructureOpen Source: Yes

Graviton

Run 500B+ parameter LLMs locally on consumer hardware with streaming quantization.

Visit Graviton

Graviton streams model layers from SSD and quantizes in-flight (FP16 to 4-bit/2-bit/1.58-bit) so models never need to fit in memory all at once. A 72B model compresses to 36 GB and runs on a 64 GB Mac. Includes speculative decoding for 2-3x throughput, dynamic sparsity, REST API for agents, and a chat UI.

FREE TIER
Completely free and open-source under Apache 2.0 license.
COMPANY
OpenGraviton
CATEGORY
InfrastructureLocal Inference
OPEN SOURCE
Yes
PRIVACY
Runs entirely locally, no data leaves your machine
TAGS
local-llmquantizationinferenceapple-siliconopen-sourceself-hosted

BEST FOR

  • Users with consumer hardware running huge LLMs locally, even 500B+.
  • Developers building private AI agents; offers local REST API & performance.
  • Privacy-focused users requiring powerful LLMs without data ever leaving device.

NOT FOR

  • Users seeking a simple, zero-setup cloud-based LLM experience.
  • Non-technical users expecting a plug-and-play chat interface without setup.
Finding similar tools…
← Back to all tools