Replicate MCP integration
Finds models and reads their input schemas, runs and cancels predictions, manages deployments and hardware, and drives fine-tuning jobs.
31actions available
Three actions you can hand over today
Every action runs live through MCP. Nothing to build, nothing to maintain.
Create model prediction
Lyro creates predictions using Replicate models to process customer data in real time. This brings AI model capabilities directly into conversations without requiring external API calls.
Search models and collections
Lyro searches the Replicate model library to find the right model for a given customer need. Agents locate specialized models instantly instead of browsing Replicate manually.
Cancel prediction
Lyro cancels long-running predictions that are no longer needed. Customers get faster responses when Lyro stops unnecessary processing mid-conversation.
How businesses use Replicate + Lyro
Each card is one request a support team gets, and the Replicate actions Lyro runs to close it.
Find the right model before committing to it
Lyro searches models, collections, and docs by keyword, pulls a candidate's details and input schema, and reads its README and author-provided example predictions to see how it is meant to be called.
Search Models and CollectionsGet Model DetailsGet Model READMERun a prediction and keep an eye on it
Lyro creates a prediction against an official model or a specific version, polls its status and output by ID, and cancels one that is still running when it is clearly going nowhere.
Create Model PredictionGet PredictionCancel PredictionPut a model into production on the right hardware
Lyro checks which hardware SKUs are available, creates a deployment with the model version and scaling parameters it needs, and reports back on an existing deployment's release configuration.
List Available HardwareCreate DeploymentGet Deployment DetailsDrive a fine-tuning run end to end
Lyro creates the destination model, launches a training job against a model version with your training data, and lists or cancels the jobs already in flight.
Create ModelCreate Training JobCancel Training
How it works
Get started in 3 steps
Connect once, then just ask. There is no workflow builder to learn and nothing to maintain — Lyro reads the Replicate actions it has and picks the ones a request needs.
- 01
Connect Replicate
Authorize the Replicate account your team already uses — one consent screen, no API keys, no mapping tables. Lyro can only do what you granted that account, and you can disconnect it at any time.
- 02
Tell your agent what you need
Describe the job the way you would hand it to a teammate. Lyro maps it to the Replicate actions that close it and chains as many as the request needs.
- 03
Watch it work
The agent runs the actions inside the conversation the customer is already in, so nobody copies data between tabs and your team can take over at any point.
Get started free
Everything else about Replicate
Setup, permissions, and the limits of what Lyro can do inside Replicate.
It depends on what you are running. Create Model Prediction targets an official model by owner and name, the version-based Create Prediction runs a specific model version ID, and the deployment-based Create Prediction only works against a Replicate Deployment, meaning a persistent instance you have already created. Pointing the deployment action at a plain model will not work.
Every action available in Replicate
All 31 actions your agent can call on Replicate, straight from the live MCP connection.
Get account information
Get authenticated account information.
Cancel prediction
Cancel a prediction that is still running.
Get model collection
Get a specific collection of models by its slug.
List model collections
List all collections of models.
Create model
Create a new Replicate model with specified owner, name, visibility, and hardware.
Create prediction
Create a prediction for a Replicate Deployment.
Create deployment
Create a new deployment with specified model, version, hardware, and scaling parameters.
Delete deployment
Delete a deployment from your account.
Get deployment details
Get deployment details by owner and name.
List deployments
List all deployments associated with the account.
Create file
Create or upload a file to Replicate.
Delete file
Delete a file by its ID.
Get file details
Get details of a file by its ID.
List files
Retrieve a paginated list of uploaded files.
Get prediction
Get the status and output of a prediction by its ID.
List available hardware
List available hardware SKUs for models and deployments.
List model examples
List example predictions for a specific model.
Get model details
Get details of a specific model by owner and name.
List public models
List public models with pagination and sorting.
Create model prediction
Create a prediction using an official Replicate model.
Get model readme
Get the README content for a model in Markdown format.
Get model version
Get a specific version of a model.
List model versions
List all versions of a specific model.
Create prediction
Create a prediction to run a model by version ID.
List all predictions
List all predictions for the authenticated user or organization with pagination.
Search models and collections
Search for models, collections, and docs using text queries (beta).
Cancel training
Cancel an ongoing training operation in Replicate.
Create training job
Create a training job for a specific model version.
List training jobs
List all training jobs for the authenticated user or organization with pagination.
Update model metadata
Update metadata for a model including description, URLs, and README.
The tools Replicate sits next to
Same connection, same setup. Pick the next one your team already uses.
Astica.ai
Pulls printed and handwritten text out of images a customer attaches, and transcribes audio clips into text the conversation can use.
DeepImage
Submits enhancement jobs, tracks them by hash, returns the finished result URL, and clears stored images once they are delivered.
Hugging Face
Vets a model's card and security scan before adoption, runs inference, commits to Hub repositories, and watches repos through webhooks.
Modelry
Lists 3D and AR embeds, orders modeling jobs, and reports completion percentage on the ones already in progress.
RunPod
Reports GPU availability and pricing, provisions clusters and serverless endpoints, and stores the registry credentials a private image needs.
Vectorshift
Runs pipelines singly or in bulk, stops an execution by run ID, and creates, inspects, or removes the chatbots built on top of them.

Ready to connect Replicate?
Authorize the account and your agent has all 31 actions from the first conversation.


