Anyscale is the company behind Ray, the open-source distributed computing framework born at UC Berkeley's RISELab. Its managed platform runs training, batch inference, and LLM serving workloads on autoscaling Ray clusters, hosted or in your own cloud (BYOC).

Explore IT tools
Filter by deployment, source model and use case.
Refine results
23 results
AnythingLLM is a free, MIT-licensed desktop and self-hostable app (by Mintplex Labs) for chatting with your own documents locally, with a separate paid team/cloud option.
Cohere provides text generation, embedding and rerank models via API for enterprise use, alongside North (an AI workplace platform) and Compass (enterprise search), with free trial keys and usage-based production pricing.
CrewAI is an MIT-licensed open-source Python framework for orchestrating multi-agent AI workflows; CrewAI Enterprise adds a hosted runtime with a free Basic tier and custom-priced governance features.
DeepInfra provides pay-as-you-go, serverless GPU inference APIs for 100+ open-source AI models covering text generation, text-to-speech, text-to-image and embeddings, without customers managing their own GPU infrastructure.
Dify (by LangGenius) is an open-source platform for building LLM apps with a visual workflow builder, RAG pipeline and agent tools; self-host the Community Edition free or use Dify Cloud on a freemium plan.
Fireworks AI provides low-latency, cost-optimized hosted inference and fine-tuning for open-source language and image models, using its own inference stack as the main differentiator versus other model-hosting APIs.
Flowise is an open-source, self-hostable LLM/agent builder acquired by Workday in 2025. Workday announced Flowise is being sunset: feature work stopped and core support ended 31 August 2026.
GPT4All is a free, open-source desktop application from Nomic AI that runs open-source large language models locally on Windows, macOS and Linux, with no data leaving the device.
LangChain is an MIT-licensed open-source framework (Python/JS) for building LLM applications and agents via LangGraph; LangSmith adds optional hosted tracing, evaluation and deployment.
LlamaIndex is an MIT-licensed open-source framework (Python/TypeScript) for RAG pipelines and data agents on top of LLMs; LlamaCloud adds hosted document parsing and extraction.
Mem is an AI-powered notes and knowledge-management app that automatically organizes notes and resurfaces related ones, with an AI chat assistant for querying your own content.
Milvus is an open-source vector database for large-scale similarity search in AI apps, self-hostable as standalone or a distributed cluster, or via managed Zilliz Cloud.
Mistral AI provides model access and developer tools through its Studio and API, alongside Le Chat and coding-oriented products.
OctoAI (formerly OctoML) was a GPU inference platform for generative AI models. NVIDIA acquired it in September 2024 and shut down OctoAI's public cloud service on October 31, 2024; octo.ai now redirects to NVIDIA's website.
Ollama is an open-source, MIT-licensed tool for running LLMs on your own hardware via a local REST API and CLI; local use is free and unlimited, with an optional paid Ollama Cloud for larger models.
Replicate lets developers call thousands of public open-source models or deploy their own via Cog, its open-source packaging tool, billed by actual compute time with no permanent free tier.
Runpod supplies GPU Pods for interactive or batch workloads and serverless endpoints for supported inference deployments.
Scale AI supplies data annotation, human-feedback (RLHF) pipelines, and model evaluation services for training and testing AI models; engagements are sold via direct sales, not self-serve pricing.
SuniAI was listed as Sunitech's AI offering, but sunitech.ch/ia now redirects to the general homepage with no distinct AI product description, so its scope needs direct vendor confirmation.
Tana is an agentic meeting platform: AI agents join video calls and take real-time actions like filing issues and updating trackers, on top of a persistent, graph-based context that also powers the original Tana Outliner notes tool.
Together AI runs an API platform for inference and fine-tuning of open-source models, plus dedicated GPU cluster rentals, positioned as an alternative to building your own model-serving infrastructure.
Weights & Biases (W&B) tracks ML experiments, manages model/dataset versions, and evaluates LLM applications via Weave, available as hosted SaaS, self-hosted Server, or an enterprise private cloud.