Describe your analytics needs in plain English. DataMCP connects to your Postgres database, maps the schema, generates a production-quality dbt project, wires up dashboards, schedules orchestration, and streams real-time data — all without leaving your coding assistant.
The tools all exist. dbt, Dagster, Metabase, Kafka — the ecosystem is mature. The problem is the assembly cost: a skilled data engineer needs one to three weeks just to wire everything together for the first time. Then another week whenever the schema changes.
DataMCP runs as an MCP server inside Claude Code, Cursor, or Codex. The developer stays in their editor. Every tool call is logged to a local knowledge base that makes every subsequent interaction smarter.
Every DataMCP session follows the same reproducible flow — from raw Postgres to analytics-ready marts, with tests, documentation, and monitoring at each stage.
Each stone adds a complete layer of functionality — independently useful, but designed to compose. A team can adopt Stone 1 today and unlock the full stack as they grow.
Every DataMCP action is stored in a local SQLite knowledge base — schema snapshots, model generations, KPI definitions, pipeline run history, query logs, Metabase chart metadata, Kafka connection configs. This is the context that makes every subsequent interaction faster and more specific.
Persisted locally in the project directory. Grows with the project. Never leaves the team's environment.
Every file DataMCP generates lives in the team's repository, under their version control. The generated dbt models follow naming conventions, include column-level documentation, apply correct materialisation strategies, and ship with data quality tests. There is no proprietary format. The output is standard dbt.
not_null + unique tests on every primary keyrelationships tests on all detected foreign keysschema.yml for every modelAll files committed to the repo. Readable, editable, owned by the team.
Every stone ships with full test coverage, typed interfaces, and CI-ready structure. The MCP protocol contract is stable across all six stones.
Each tool is a typed Python function registered with the MCP server. Tools are organised by capability layer and can be composed by the AI in any sequence to complete complex data engineering tasks.
DataMCP is a single MCP server process with a typed tool registry. All tools share a common knowledge base client and a database connection pool. The MCP protocol contract is stable — adding new tools never changes the interface for existing ones.
DataMCP is built on the established open-source data engineering ecosystem. Every tool it generates uses standards that any data engineer already knows.
Each stone was scoped to be independently useful — a team gets real value from Stone 1 alone on day one. The full stack unlocks progressively as stones are added.
The full source — MCP server, 20 tools, 6 stones, Jinja2 templates, SQLite knowledge base, Dagster integration, Kafka connectors, and 232 passing tests — is available on request. Built as a portfolio project demonstrating production-grade AI data engineering.
✉ Request Repository Access