Serving models, routing between providers, capping spend, and fine-tuning when prompting genuinely is not enough.
The gateway tools pay for themselves the first month you would otherwise have found out about overspend from the invoice.
Moving fastest this week: litellm (+2.3k), langfuse (+1.7k), promptfoo (+873).
Mature enough to put in production this quarter.
Worth a timeboxed spike before you bet on it.
AI model testing and evaluation framework
Instead of: manual model testing and evaluation
Open-source LLM evaluation and monitoring tool
Instead of: Manual model evaluation and monitoring
Mixture-of-Models router for LLM inference
Instead of: Manual model routing or paid AI gateways
LLM serving kernel library
Instead of: custom LLM serving code
Open-source model management platform
Instead of: Manual model tracking and logging
Rust framework for modular, scalable LLM applications
Instead of: Hand‑rolled Python glue code or costly LLM‑ops SaaS
No-code GUI for fine-tuning LLMs
Instead of: Manual LLM fine-tuning or paid GUI tools
Swift library for Gemma 4 model inference
Instead of: Cloud-based model inference services
Self-hosted AI-powered wiki with C# backend and TypeScript frontend
Instead of: Notion, Confluence, or GitBook plus a separate RAG pipeline
AI model optimization and deployment toolkit
Instead of: Manual model optimization and deployment scripts
Open-source LLM fine-tuning framework
Instead of: Manual model tuning or paid services
C99 inference engine for 2.78 T‑parameter Kimi K3, runs on a single CPU with ~8 GB RAM
Instead of: cloud LLM API calls or heavyweight GPU inference stacks
Rust WebGPU inference engine
Instead of: Paid inference services or Python-based engines
Parameter-efficient fine-tuning for large language models
Instead of: Manual fine-tuning or paid model optimization services
AI workflow orchestration tool
Instead of: manual scripting or paid workflow tools
Open-source LLM API endpoint
Instead of: Manual LLM deployment and management
Python framework for serving AI models
Instead of: Custom model serving code
Python framework for building ML workflows
Instead of: Manual scripting or paid ML platforms
Toolkit for compressing and serving large language models
Instead of: Manual model optimization and deployment scripts
AI model serving platform for Kubernetes
Instead of: custom model deployment scripts
Fine-tune and deploy open LLMs
Instead of: Manual model tuning and deployment
Open-source voice generation model
Instead of: Paid voice generation APIs
AI gateway for routing to LLMs
Instead of: manual LLM integration
Multi-LoRA inference server for fine-tuned LLMs
Instead of: Manual model deployment scripts
Real software; just not where a small team's next hundred hours should go.
AI engineering platform for LLMs
Instead of: manual LLM evals and metrics
AI model monitoring and evaluation tool
Instead of: manual model evaluation and monitoring
Open-source AlphaEvolve implementation
Instead of: Manual evolutionary algorithm coding
A single interface for fine-tuning a hundred-plus open models without writing training code.
Instead of: A machine-learning contractor's setup time.
AI agent development framework
Instead of: custom AI tooling
AI proxy server for model routing and observability
Instead of: custom model routing logic
Open LLM devops platform
Instead of: Manual model management
AI model building and evaluation toolkit
Instead of: Manual model evals and custom scripts
Pre-trained LLMs with deployment scripts
Instead of: Manual LLM training and deployment
Large language model infrastructure tools
Instead of: Manual LLM deployment scripts
Distributed inference serving framework
Instead of: Custom model serving infrastructure
Open-source LLMOps platform
Instead of: Manual model management
Mistral model inference library
Instead of: Custom model serving code
Ray is a distributed compute engine for scaling ML training and inference, built for large teams with dedicated infrastructure staff.
Instead of: Manual scripting with Python multiprocessing or cloud-based ML platforms like SageMaker or Comet.
Local large language model serving
Instead of: Cloud-based LLM services
AI agent for automating dev tasks
Instead of: manual dev task automation
JavaScript tool for creating LLM datasets
Instead of: manual dataset creation
Tell us what you sell in one sentence and we will hand you three specific moves — priced per month, with the arithmetic shown. Free, no signup.
Give me three moves