beam
beam
Discover
PulseActivityAnalyticsBest forMapOrgs
Niches
AgentsMCPRAGCoding AssistantsInference & ServingVector DBs
Personal
WatchlistCompareWeekly report
?
Sign in
magic link · no password
LIVE──────── · ──:──:── UTCabout
beam
beam
Discover
PulseActivityAnalyticsBest forMapOrgs
Niches
AgentsMCPRAGCoding AssistantsInference & ServingVector DBs
Personal
WatchlistCompareWeekly report
?
Sign in
magic link · no password
All niches
Nicheinference

Inference & Serving

[TOOLS_TRACKED]32
[ACCELERATING]0
[DYING]0
[AVG_VELOCITY]0.56/10
See our pick → Best Inference & Serving
[ACCELERATING]

Accelerating

0

No accelerating tools right now.

[STABLE]

Stable

12
Tool
Velocity
Trend 30d
Δ 7d
Stars
Class
  • ray-project/rayRay is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
    2.27↑ +7144kStable
  • google-ai-edge/mediapipeCross-platform, customizable ML solutions for live and streaming media.
    2.01↑ +6537kStable
  • NVIDIA/nvcfPlatform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.
    1.93↑ +7221Stable
  • tetherto/qvacOpen-source local AI SDK - run AI on-device with no cloud, no API keys. Supports GGUF, RAG, image, music, and video generation, speech-to-text, P2P inference, and more. Cross-platform: Linux, macOS, Windows, Android, iOS.
    1.54↑ +19624Stable
  • gpustack/gpustackA GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
    1.22↑ +395.7kStable
  • Avarok-Cybersecurity/atlasPure Rust Inference Engine
    1.05↑ +18700Stable
  • OpenRLHF/OpenRLHFAn Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
    0.87↑ +2410kStable
  • xorbitsai/inferenceSwap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
    0.77↑ +189.6kStable
  • dstackai/dstackA unified orchestration layer for heterogeneous AI compute. It standardizes how to manage compute and run training and inference on GPU clouds, Kubernetes, VMs, or bare-metal clusters.
    0.77↑ +92.3kStable
  • Lightning-AI/litgpt20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
    0.66↑ +1214kStable
  • bentoml/BentoMLThe easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
    0.66↑ +118.9kStable
  • mozilla-ai/any-llmCommunicate with an LLM provider using a single interface
    0.65↑ +82.2kStable
LIVE──────── · ──:──:── UTCabout