AI-Hypercomputer/JetStream
Google's throughput and memory optimized inference engine for running LLMs on TPUs and GPUs.

Not currently ranked — collecting fresh signals.
star history
JetStream is an LLM inference engine designed for high throughput and memory efficiency on XLA-based accelerators, primarily TPUs with GPU support coming. It provides reference implementations for both Jax and Pytorch model execution, enabling efficient serving of models like Llama, Gemma, and GPT variants on Google Cloud TPU infrastructure.
Frequently asked
- What is AI-Hypercomputer/JetStream?
- Google's throughput and memory optimized inference engine for running LLMs on TPUs and GPUs.
- Is JetStream open source?
- Yes — AI-Hypercomputer/JetStream is open source, released under the Apache-2.0 license.
- What language is JetStream written in?
- AI-Hypercomputer/JetStream is primarily written in Python.
- How popular is JetStream?
- AI-Hypercomputer/JetStream has 448 stars on GitHub.
- Where can I find JetStream?
- AI-Hypercomputer/JetStream is on GitHub at https://github.com/AI-Hypercomputer/JetStream.