Appendix C — Library and application responsibilities

Libraries provide operations such as searching an index or validating returned fields. The surrounding software still needs rules for access, data updates, acceptable results, and failure handling. The table separates those tasks and identifies records that help check their behavior. The reference architecture in Section 12.1 shows where each library sits in one system.

Table C.1: Each library provides specific operations while application code supplies task rules and checks.
Software or protocol Representative use What remains the application’s responsibility Primary operational evidence
OpenAI-compatible client / Nebius endpoint Send versioned chat or completion requests and read usage Task specification, prompt quality, provider choice, privacy decision, or release threshold Request ID and input/prompt versions, model/version, token counts, latency and response status
Pydantic Validate structured model output against typed fields and constraints Semantic correctness, missing-evidence policy, or side-effect authorization Validation result, rejected payload, repair or escalation path
PyPDFLoader / document loader Extract page text and preserve document/page identity Source reliability, parsing correctness, access policy, or deletion lifecycle Document hash, parser version, page count, extraction warnings
Recursive text splitter Create bounded overlapping chunks The right semantic unit, overlap policy, or retrieval quality Source document/page and chunk IDs, split positions, token counts and splitter configuration
BGE embedding model Map queries and chunks into a retrieval vector space Domain fit, access filtering, index freshness, or semantic truth Model revision, vector dimension, normalization, benchmark slice
FAISS Persist and search a vector index Document authorization, metadata correctness, freshness, or relevance Index version, build manifest, candidate IDs and distances
BM25 / lexical index Retrieve exact names, identifiers, and rare terms Synonym coverage or grounded answer generation Analyzer configuration, term scores, corpus/index version
BGE reranker Reorder a candidate set with query-document interaction Recall lost before reranking or final answer faithfulness Candidate set, reranker revision, before/after ranks, latency
Ragas Compute configured RAG measures, including model-assisted metrics Metric validity for the domain, judge calibration, or the final release decision Pinned package version, judge model, metric inputs, raw result
LangGraph Define graph state and routing, with optional configured checkpointing and resume Business policy, tool permissions, state schema quality, or completion criteria Graph version, node transitions, checkpoint ID, final state
FastMCP / MCP server Expose typed tools, resources, and prompts through MCP messages between a host’s client and a server Host consent, credential scope, server trust, or downstream authorization Protocol/version, capability negotiation, principal, arguments, result provenance
Container or code sandbox Constrain generated code execution and collect outputs Whether code execution is necessary, safe inputs, or acceptable side effects Image hash, resource limits, network/filesystem policy, exit status
Experiment tracker Record dataset, prompt, model, code, metrics, and artifacts The task definition, label quality, causal interpretation, or release decision Immutable run ID, linked data, prompt, model, code and result versions, comparison and approval record
Telemetry and tracing Observe latency, errors, token/tool use, and state transitions The meaning of success, which private data may be logged, or retention policy Trace ID, redaction status, sampled events, metric definitions