ashhart/TensorFold
TensorFold is an LLM serving system that runs text models on Apple Silicon (MLX) and NVIDIA GPUs behind an OpenAI-compatible API.

TensorFold serves large language models through an OpenAI-compatible HTTP endpoint, supporting multiple model families with their own kernels and draft verification for speculative decoding. It targets Apple Silicon via MLX and NVIDIA GPUs via CUDA, with support for techniques such as multi-token prediction heads, DFlash2 drafting, and context copies. The CLI provides commands to list models, inspect configurations, pull checkpoints, and start a server.
Frequently asked
- What is ashhart/TensorFold?
- TensorFold is an LLM serving system that runs text models on Apple Silicon (MLX) and NVIDIA GPUs behind an OpenAI-compatible API.
- Is TensorFold open source?
- Yes — ashhart/TensorFold is open source, released under the MIT license.
- What language is TensorFold written in?
- ashhart/TensorFold is primarily written in Python.
- How popular is TensorFold?
- ashhart/TensorFold has 558 stars on GitHub.
- Where can I find TensorFold?
- ashhart/TensorFold is on GitHub at https://github.com/ashhart/TensorFold.