Problem
The GitHub issues ingestion pipeline chunks issue content up to 1500 characters, but truncates every chunk to 1000 characters before sending it to the TEI embeddings service:
docs-agent-mcp/pipelines/issues-pipeline.py#L481 — chunk_size: int = 1500
docs-agent-mcp/pipelines/issues-pipeline.py#L301 — max_tei_chars = 1000, applied as r["content_text"][:max_tei_chars]
So for any chunk longer than 1000 chars, the last ~500 characters are stored in Milvus and shown to users, but were never embedded — that text is invisible to semantic search.
Problem
The GitHub issues ingestion pipeline chunks issue content up to 1500 characters, but truncates every chunk to 1000 characters before sending it to the TEI embeddings service:
docs-agent-mcp/pipelines/issues-pipeline.py#L481—chunk_size: int = 1500docs-agent-mcp/pipelines/issues-pipeline.py#L301—max_tei_chars = 1000, applied asr["content_text"][:max_tei_chars]So for any chunk longer than 1000 chars, the last ~500 characters are stored in Milvus and shown to users, but were never embedded — that text is invisible to semantic search.