Projects02 AI / Voice
Multi-Modal AI Agent Platform
A voice and webhook platform where ElevenLabs, an LLM, MCP, and n8n turn a spoken or event-driven request into an integration run.
Problem
Voice, chat, and backend automations usually sit in separate tools, so a spoken request or an inbound webhook cannot drive the same workflow.
Solution
Built a multi-modal agent platform: ElevenLabs handles voice, webhooks bring events in, an LLM chooses the next step, and MCP plus n8n run the integrations. The same agent path serves speech and machine events.
Architecture
- Voice
- Webhook
- LLM
- MCP
- n8n
- Integrations
Technology
- n8n
- ElevenLabs
- Webhooks
- LLM
- MCP
Engineering decisions
- The same intent from speech and from a webhook, without two different agents.
- Voice transcripts that are good enough to act on, and a spoken reply that is short enough to hear.
- Webhooks that authenticate the caller before any tool runs.
- MCP and n8n staying the execution layer so the model does not talk directly to every API.
Automation
A completed intent triggers the matching n8n workflow. The model does not hold the integration credentials.
Reliability
A bad transcript or a failed webhook is acknowledged and retried; it does not fire an integration with an empty payload.
Security
Webhook secrets and ElevenLabs keys stay off the client. MCP tools are allowlisted per workflow.
AI layer
The LLM turns speech or an event into a structured intent. MCP exposes tools. n8n performs the side effect after the intent is accepted.