Skip to content
Srikanta Sahu

Projects02 AI / Voice

Multi-Modal AI Agent Platform

A voice and webhook platform where ElevenLabs, an LLM, MCP, and n8n turn a spoken or event-driven request into an integration run.

Problem

Voice, chat, and backend automations usually sit in separate tools, so a spoken request or an inbound webhook cannot drive the same workflow.

Solution

Built a multi-modal agent platform: ElevenLabs handles voice, webhooks bring events in, an LLM chooses the next step, and MCP plus n8n run the integrations. The same agent path serves speech and machine events.

Architecture

  1. Voice
  2. Webhook
  3. LLM
  4. MCP
  5. n8n
  6. Integrations

Technology

  • n8n
  • ElevenLabs
  • Webhooks
  • LLM
  • MCP

Engineering decisions

  • The same intent from speech and from a webhook, without two different agents.
  • Voice transcripts that are good enough to act on, and a spoken reply that is short enough to hear.
  • Webhooks that authenticate the caller before any tool runs.
  • MCP and n8n staying the execution layer so the model does not talk directly to every API.

Automation

A completed intent triggers the matching n8n workflow. The model does not hold the integration credentials.

Reliability

A bad transcript or a failed webhook is acknowledged and retried; it does not fire an integration with an empty payload.

Security

Webhook secrets and ElevenLabs keys stay off the client. MCP tools are allowlisted per workflow.

AI layer

The LLM turns speech or an event into a structured intent. MCP exposes tools. n8n performs the side effect after the intent is accepted.