Back to blog
How to Build a Knowledge Graph From Meeting Notes

This blog is written by AI for SEO

How to Build a Knowledge Graph From Meeting Notes

HelixDB15 min read

Every engineering team eventually dumps hundreds of Granola, Otter, or Zoom transcripts into a vector database, sets up top-k cosine similarity, and expects an AI agent to answer questions about past decisions. Within forty-eight hours, the cracks appear. The agent confuses which engineer agreed to take on an API refactor, conflates two separate roadmap debates from different months, and cites discarded proposals as active architecture plans.

Meeting transcripts are not static documents like technical specs or help-desk articles. They are conversational logs where context drifts fast, speakers interrupt one another, decisions overwrite earlier consensus, and commitments bind specific people to specific tasks. Blindly vectorizing raw text loses every one of those structural relationships. To make meeting records actually usable by autonomous agents and internal copilots, you have to structure the conversational data properly.

This guide shows you how to build a knowledge graph from meeting notes, from raw ingestion to graph traversal. You will learn the exact schema of entities and edges required for multi-speaker transcripts, how to extract and disambiguate them across time, and how to query the resulting graph so your agents can retrieve what actually happened.

Why Does Embedding a Transcript as One Document Break Everything?

Dumping a 45-minute audio transcript directly into an embedding model crushes structural nuance into a single high-dimensional average. An hour-long product sync easily spans 8,000 words. Embedding that entire block into a single 1536-dimensional vector guarantees that specific action items, contentious objections, and decisions get washed out by conversational filler.

Even standard chunking strategies (like splitting every 500 tokens with a 50-token overlap) fail on conversational data. A single topic might begin at minute four, pause for a tangent about CI/CD pipelines, and resolve at minute twenty-two with a concrete commitment. When you chunk by token length, the sentence where a tech lead says "Okay, let's ship the migration on Friday" gets divorced from the preceding context that specified which database migration they were discussing. The chunk retains the semantic vector for scheduling a launch, but loses the entity the launch applies to.

Standard vector search treats every chunk as an isolated semantic island. When an agent queries "What did Sarah commit to for the billing refactor?", cosine similarity searches for chunks that match that semantic phrasing. If Sarah simply said "I will handle the schema migration by Thursday", the word "billing" never appears in her utterance. Vector search misses the chunk entirely or surfaces a chunk from an unrelated billing discussion six months earlier. To preserve causality, accountability, and sequence, you have to model the meeting as an interconnected graph rather than flat text.

What Are the Actual Entities in a Meeting Transcript?

Before writing an extraction prompt, you need a strict domain model. Most failed graph implementations collapse because the engineering team lets an LLM invent arbitrary node labels like "DiscussionPoint" or "Thoughts". A functional meeting knowledge graph requires a defined set of node types.

First is the Meeting node. This captures the session metadata: date, duration, calendar title, and unique transcript identifier. Second is the Participant node. This represents the real humans in your organization, storing their canonical name, email, and company role.

Third is the Utterance node. This is the atomic conversational unit, capturing raw text spoken by a specific person at a specific millisecond timestamp. Fourth is the Topic node, which clusters several contiguous utterances around a unified subject (such as "Auth Migration" or "Q3 Hiring Budget").

Fifth and sixth are the critical operational entities: Decisions and Action Items. A Decision node represents an agreement reached during the conversation, containing the explicit outcome, the rationale, and its current status. An Action Item node represents an assigned obligation, containing the task description, the assigned owner, and the deadline. Finally, create a Concept node for technical components, tools, or projects referenced during discussion, such as "PostgreSQL", "Auth0", or "Billing Engine".

Treating these nodes as distinct types transforms meeting data from conversational noise into a structured AI agent memory architecture. Your agent no longer has to guess whether a sentence was a brainstormed idea or an executed decision: the graph schema enforces the distinction.

What Edges Connect Those Entities?

Nodes hold identity, but edges hold meaning. A knowledge graph built from transcripts succeeds or fails on the precision of its edge relationships. Without typed edges, you cannot traverse from a high-level corporate goal down to the engineer who agreed to complete it.

Connect participants to meetings using ATTENDED edges, which can carry properties like role in meeting (such as organizer or attendee). Connect each Utterance to its speaker using a SPOKE edge, and link sequential utterances together using a NEXT edge to preserve chronological conversational flow.

Link Utterance nodes to their overarching Topic node using a BELONGS_TO edge. When an utterance introduces or ratifies an agreement, draw a LED_TO edge directly from the Utterance to the Decision node. This gives you an audit trail: you can trace any decision back to the exact timestamp and context where it was formulated.

For operational tracking, draw an ASSIGNED_TO edge from an Action Item to a Participant, and an ORIGINATED_IN edge from the Action Item back to the Meeting node. When an Action Item stems directly from an agreed architectural change, draw an IMPLEMENTS edge from the Action Item to the Decision.

Finally, connect Decisions, Action Items, and Topics to relevant technical systems using a REFERENCES edge targeting Concept nodes. When your system needs to determine why a service was deprecated, it can traverse from the Concept node backward across REFERENCES edges to locate the Decision, the parent Meeting, and the engineering debate that produced it.

How Do You Run the Extraction Pass?

Never attempt to extract your entire graph in a single massive prompt over an hour-long transcript. Passing a lengthy conversational log into an LLM and demanding nodes, edges, properties, and JSON schema formatting in one shot often leads to hallucinations and missed entities.

Structure your extraction as a two-pass pipeline. Pass one is segmentation and topic boundary detection. Run a lightweight model across the raw transcript to identify topic transitions. Prompt the model to group contiguous utterances into discrete segments marked by start and end timestamps. This turns a sprawling conversation into three to eight bounded, coherent topical units.

Pass two is typed entity and relationship extraction executed independently across each segment. Feed the segment text alongside the known participant list from the calendar event. Instruct the model to return a structured JSON payload conforming strictly to your schema. Use JSON schema enforcement or tool calling to guarantee valid typing.

Here is an example prompt structure for pass two:

{
  "meeting_id": "mtg_88492",
  "segment_index": 2,
  "topic": "Database Indexing Strategy",
  "decisions": [
    {
      "id": "dec_01",
      "text": "Move the billing schema migration behind a feature flag",
      "rationale": "Lets us roll back without a second deploy window",
      "source_utterance_indices": [14, 15]
    }
  ],
  "action_items": [
    {
      "id": "act_01",
      "assignee": "Elena Rostova",
      "task": "Benchmark the migration path under 10k writes/sec",
      "deadline": "2026-04-15",
      "source_utterance_indices": [18]
    }
  ],
  "concepts_referenced": ["Billing Schema", "Feature Flags"]
}

By isolating extraction to bounded segments, the LLM retains enough attention to catch offhand commitments that monolithic prompts skip entirely.

How Do You Chunk the Transcript Before Embedding?

Extracting symbolic entities and relationships solves structural navigation, but your agent still needs semantic vector search to find relevant passages when an engineer asks natural language questions. You have to embed the transcript text, but chunking transcripts requires a dialogue-aware strategy.

Never chunk transcripts using arbitrary character counts or fixed token windows. A fixed window splits mid-sentence and severs the speaker attribution from the spoken thought. Instead, chunk along utterance and speaker turn boundaries.

Combine multiple contiguous utterances within the same topic segment until you reach a semantic target, typically 250 to 500 tokens. If a speaker talks continuously for 600 tokens, keep that utterance intact as a single chunk. When speakers exchange brief, rapid comments during a debate, group those turns together into a single dialogue block.

Every chunk must carry prepended contextual metadata before passing through your embedding model. Prepend the meeting title, the date, the current topic, and the speakers active in that chunk directly to the input text. For example:

[Meeting: Core Infra Sync | Date: 2026-03-12 | Topic: Storage Backend | Speakers: Alex, Dave]

Prepending this context anchors the embedding vector in the semantic space of the overarching project. Without this header, a chunk containing "I think we should switch from local storage to remote object storage" yields an isolated semantic vector. With the header, the embedding explicitly aligns with queries about infrastructural decisions for the Core Infra team. Learn more about structuring your ingestion stages in our guide on how to build a GraphRAG pipeline.

How Do You Handle Entity Disambiguation Across Meetings?

Transcripts contain messy, incomplete references. In meeting one, a speaker is called "Dave". In meeting two, he is listed by his calendar email "david.k@company.com". In meeting three, colleagues refer to him as "Dave from platform". If your extraction pipeline creates three distinct Participant nodes, your knowledge graph fragments immediately.

Entity disambiguation must run before inserting nodes into the graph. For people, maintain a canonical resolution table seeded by your company directory or Google Workspace. When an extraction pass surfaces a speaker named "Dave", map the string through three checks: first, exact match against known meeting attendees via calendar metadata; second, alias match against known nicknames; third, email match. If an extraction yields an unresolvable name, compute a string distance or vector similarity against the organization's verified user list before creating a new node.

Apply the same rigor to Concept nodes. Engineers constantly alternate between acronyms, internal code names, and informal shorthand. One engineer says "the ingestion service", another says "Project Falcon", and a third says "the Rust pipeline". Maintain an alias array on your Concept nodes:

{
  "canonical_name": "Falcon Ingestion Pipeline",
  "aliases": ["Falcon", "ingestion service", "rust pipeline", "stream worker"],
  "type": "InternalService"
}

When extracting concepts from a new meeting, resolve names against existing canonical entities using token normalization and vector similarity over alias lists. Set a similarity threshold you have actually calibrated against your own alias list, and merge above it. Do not create duplicate nodes for identical architectural systems. The same problem shows up across every internal source, not just transcripts: one person is David Miller in the directory, d.miller on GitHub and Dave in Slack, which is why semantic search over internal documents is not enough on its own.

How Do You Store Vectors and Graph Edges in the Same Write?

The common setup keeps edges in a graph database and embeddings in a separate vector store, with glue code holding the two in sync. That glue is where transcript pipelines rot: an edge write lands, the matching vector insert fails, and your agent retrieves a Decision node whose text it cannot rank.

HelixDB is an open-source graph-vector database written in Rust. Vectors are a property on nodes and edges, so the relational edge, the embedding, and the metadata all persist in one write. There is no separate query language to learn and nothing compiled ahead of time: the SDK builders run inside your own code and serialise to JSON, which the database turns into the query it executes. Writes go to POST /v2/query, port 6969 on a local instance.

Start with the preconditions, because they are what make a nightly re-run safe. Give Participant and Meeting a unique-equality index on the stable id you chose during extraction:

import { g, writeBatch, IndexSpec } from "@helix-db/helix-db";

const setup = writeBatch()
  .varAs("participantIdx", g().createIndexIfNotExists(IndexSpec.nodeUniqueEquality("Participant", "email")))
  .varAs("meetingIdx", g().createIndexIfNotExists(IndexSpec.nodeUniqueEquality("Meeting", "transcriptId")))
  .returning(["participantIdx", "meetingIdx"]);

Index creation returns before the backfill finishes, so a new index is not usable the instant the call is accepted. Then write the meeting, the decision and the edges in one batch:

import { g, writeBatch, NodeRef, SourcePredicate, BatchCondition } from "@helix-db/helix-db";

const query = writeBatch()
  .varAs("meeting", g().addN("Meeting", { transcriptId: "mtg_101", title: "Architecture Review: Storage Engine" }))
  .varAs("decision", g().addN("Decision", {
    summary: "Store vectors and graph edges in one engine",
    status: "Accepted",
    embedding: [0.014, -0.042, 0.089],
  }))
  .varAs("speaker", g().nWithLabelWhere("Participant", SourcePredicate.eq("email", "dana@northwind.example")).limit(1))
  .varAs("produced", g().n(NodeRef.var("meeting")).addE("PRODUCED_DECISION", NodeRef.var("decision"), { confidence: 0.96 }).count())
  .varAsIf("ledTo", BatchCondition.varNotEmpty("speaker"),
    g().n(NodeRef.var("speaker")).addE("LED_TO", NodeRef.var("decision"), { at: "2026-03-12" }).count())
  .returning(["meeting", "decision", "speaker", "produced", "ledTo"]);

Three things in that batch are doing specific work. The .limit(1) on the speaker lookup plus the unique index is the precondition: without it, two Participant rows for the same person produce two LED_TO edges. The varAsIf gate is the other half, because a lookup that finds nobody hands addE an empty source stream, which succeeds and creates no edge. And returning every binding is how the caller tells an attached decision from an orphaned one.

That matters more than it looks, because the docs are explicit about what atomicity does and does not buy you: "All entries in the write batch commit or roll back together. Reads inside the batch can refer to earlier mutations through named variables. Do not retry a write automatically unless the full request is safe to replay." All-or-nothing means the steps land together. It does not mean your intent was satisfied, and committing a Decision with no speaker attached is a perfectly valid outcome of the batch you wrote. This is ordinary modelling in any database where a write starts with a traversal, and it is the thing that bites on Monday when you re-run the pipeline over the weekend's transcripts.

One more modelling point, and it is the one teams get wrong. The same two engineers will attend forty calls together. In most graph tooling the standard upsert idiom matches an existing relationship instead of adding one, so teams modelling repeated events end up with a single edge where they expected thousands. In HelixDB addE is additive: it adds an edge every time it is called, and an edge equality index is a lookup index rather than a uniqueness constraint, so replacing an edge is an explicit drop by id. Forty ATTENDED edges between the same pair with forty different timestamps is the correct model for "how often do these two actually work together", and it survives ingestion.

Because the edge and the embedding land in the same write, there is no window in which the graph and the index disagree about a decision that exists. The same shape scales to whatever cadence your meetings actually run at, and if you are already running a transcript graph on an in-memory engine, moving it off Memgraph maps the data model across unchanged.

What Does a Query Against This Graph Actually Look Like?

Once your meeting transcripts are structured into a unified graph-vector store, your AI agent can run hybrid GraphRAG queries that combine semantic search with relationship traversal in one pass.

Consider an agent responding to an engineer asking: "What did Sarah decide about database migrations during our March architecture meetings, and who is assigned to implement it?"

A pure vector database would search for the string "Sarah database migration March" and return isolated chunks, requiring the agent to synthesize disconnected text. In HelixDB the traversal runs first and the ranking runs inside it.

Get the order right, because it is the whole point. The documented execution order is graph traversal, then exact candidate membership, then vector ranking, then top k. You start from the Participant, walk ATTENDED to her meetings and PRODUCED_DECISION to the decisions those meetings produced, and only then rank that candidate set by similarity to the query embedding. The traversal membership is authoritative: a result outside the candidate set cannot come back. That is the opposite of searching globally and filtering afterwards, where the decisions you did not want have already consumed the source top k and you are left with fewer eligible rows than you asked for.

import { g, readBatch, defineParams, param, SourcePredicate } from "@helix-db/helix-db";

const params = defineParams({
  email: param.string(),
  query_vector: param.array(param.f32()),
  limit: param.i64(),
});

const query = readBatch()
  .varAs("decisions", g()
    .nWithLabelWhere("Participant", SourcePredicate.eq("email", params.email))
    .out("ATTENDED")
    .out("PRODUCED_DECISION")
    .vectorSearchWith("Decision", "embedding", params.query_vector, params.limit)
    .valueMap(["$id", "summary", "$distance"]))
  .returning(["decisions"]);

const request = query.toQueryRequest(
  params,
  { email: "dana@northwind.example", query_vector: [0.014, -0.042, 0.089], limit: 5n },
  { queryName: "decisions_from_attended_meetings" },
);

The honest caveat: exact membership does not mean the engine compares every candidate embedding one by one. Approximate structures still do the ranking; what is guaranteed is that the output is validated against the traversal set.

The keyword half chains the same way. textSearchWith takes the same four arguments as its vector twin, carries "$score" in the valueMap where the vector form carries "$distance", and runs the identical order. So a hybrid query over a transcript graph is scoped once and both halves respect the scope, which is what you want when an engineer half-remembers a phrase and half-remembers who said it.

Two limits worth knowing before you build on this. The server caps unrestricted vector search at 800 effective results, and a traversal-scoped search applies that ceiling after candidate intersection: it rejects when min(k, unique candidates) exceeds 800 rather than silently clamping your request. A candidate stream over 1,000,000 unique entities is a query error. The reject-rather-than-clamp behaviour is the useful half, because a silent clamp is exactly what makes a short result set inexplicable at three in the morning.

Because HelixDB supports vector search on graph edges, the traversal and vector comparison happen within the same engine scope. The agent receives a structured, deterministic subgraph containing the decision, the rationale, the meeting timestamp, and the assigned engineer.

The resulting payload fed into the LLM context window is not a jumble of conversational snippets. It is a precise relational fact: "On March 12, Sarah decided to move the billing schema migration behind a feature flag, creating Action Item #42 assigned to Elena Rostova with a deadline of April 15." The model still has to write the sentence, but it is no longer inferring who owned what from adjacent text.

Conclusion

Dumping transcripts into an isolated vector database gives your team the illusion of persistent recall while delivering brittle, untrustworthy answers. High-fidelity agent memory requires structural relationships, speaker attributions, and semantic embeddings working together.

When you build a knowledge graph from meeting notes, you turn scattered spoken words into queryable institutional memory. The extraction pass and the schema are yours either way, and most of this guide works whatever you store it in. What one engine buys you is the write that cannot half-land and the retrieval that cannot return a decision from a meeting she never attended.

The writing-data guide has the full mutation surface if you want to go further than the batch above, and the same modelling carries straight over to giving an agent persistent memory in one database. Run helix init and helix start dev to get a local instance on port 6969, and star HelixDB on GitHub if the transcript graph earns its place in your stack.

Build with HelixDB

Give your coding agent the setup prompt, or sign up and deploy a database.

Sign up