Ollama MCP integration
Route a request to a self-hosted model, reuse OpenAI-format endpoints against it, and report which models the host currently has installed.
8actions available
Three actions you can hand over today
Every action runs live through MCP. Nothing to build, nothing to maintain.
Chat with ollama model
Lyro sends a chat message with full conversation history to an Ollama model and gets back a response. The customer receives a contextually aware reply that considers prior messages.
Show model information
Lyro retrieves comprehensive details about an Ollama model including capabilities and constraints. The team confirms the model is suitable for the customer's use case.
List models
Lyro lists all available Ollama models and their specifications. The team selects the model that best fits performance and latency requirements.
How businesses use Ollama + Lyro
Each card is one request a support team gets, and the Ollama actions Lyro runs to close it.
Hold a multi-turn exchange with a self-hosted model
Lyro passes the conversation history to an Ollama model and gets a reply that accounts for it, so a step handled by your own model still follows the thread rather than starting cold.
Chat With Ollama ModelList ModelsShow Model InformationGenerate text on infrastructure you control
Lyro sends a single prompt for completion, with raw mode available when the prompt template has to be bypassed, so sensitive content is drafted without leaving your own host.
Generate Text With OllamaShow Model InformationGet Ollama VersionReuse tooling that already speaks the OpenAI format
Lyro calls the OpenAI-compatible chat and completion endpoints against local models, so prompts and integrations written for that API keep working against self-hosted inference.
OpenAI-Compatible Chat CompletionOpenAI-Compatible Text CompletionList Models (OpenAI Compatible)Check what the host can actually run
Lyro lists the installed models with size, format, and last modified date, reads one model's parameters, template, and license, and reports the Ollama version behind them.
List ModelsShow Model InformationGet Ollama Version
How it works
Get started in 3 steps
Connect once, then just ask. There is no workflow builder to learn and nothing to maintain — Lyro reads the Ollama actions it has and picks the ones a request needs.
- 01
Connect Ollama
Authorize the Ollama account your team already uses — one consent screen, no API keys, no mapping tables. Lyro can only do what you granted that account, and you can disconnect it at any time.
- 02
Tell your agent what you need
Describe the job the way you would hand it to a teammate. Lyro maps it to the Ollama actions that close it and chains as many as the request needs.
- 03
Watch it work
The agent runs the actions inside the conversation the customer is already in, so nobody copies data between tabs and your team can take over at any point.
Get started free
Everything else about Ollama
Setup, permissions, and the limits of what Lyro can do inside Ollama.
It bypasses the model's prompt templating so the prompt is sent exactly as written, which is useful for debugging or custom formatting. The trade-off is that raw mode does not return a context, so a follow-up cannot continue from it - multi-turn exchanges belong in Chat with Ollama model instead.
Every action available in Ollama
All 8 actions your agent can call on Ollama, straight from the live MCP connection.
Chat with ollama model
Send a chat message with conversation history to Ollama.
Generate text with ollama
Generate text responses from Ollama models with optional raw mode.
List models
List all available Ollama models and their details.
Openai-compatible chat completion
Create OpenAI-compatible chat completions using Ollama models.
Openai-compatible text completion
Create OpenAI-compatible text completions using Ollama models.
List models (openai compatible)
List available models using OpenAI-compatible API format.
Show model information
Show comprehensive information about an Ollama model.
Get ollama version
Get the version of Ollama running locally.
The tools Ollama sits next to
Same connection, same setup. Pick the next one your team already uses.
Chatbotkit
Pulls bot transcripts and usage figures, clones blueprints into new bots, resyncs training datasets, and wires up messaging channels.
Fal.ai
Searches model endpoints, estimates what a run will cost, follows a queued request through its logs, and returns or cancels it on request.
Kieai
Submits image, video, and music generations, polls the tasks they create, edits existing audio tracks, and reports remaining account credits.
Pinecone
Runs similarity searches and reranks results, upserts records with generated embeddings, provisions indexes and namespaces, and runs backups and imports.
Studio By Ai21 Labs
Creates and edits AI21 assistants, validates the Python behind their plans, indexes site content, and runs an assistant on request.
Zep
Writes thread messages into a user graph, retrieves the context relevant to the conversation, searches it semantically, and deletes data on request.

Ready to connect Ollama?
Authorize the account and your agent has all 8 actions from the first conversation.


