Appendix C — Library and application responsibilities
Libraries provide operations such as searching an index or validating returned fields. The surrounding software still needs rules for access, data updates, acceptable results, and failure handling. The table separates those tasks and identifies records that help check their behavior. The reference architecture in Section 12.1 shows where each library sits in one system.
| Software or protocol | Representative use | What remains the application’s responsibility | Primary operational evidence |
|---|---|---|---|
| OpenAI-compatible client / Nebius endpoint | Send versioned chat or completion requests and read usage | Task specification, prompt quality, provider choice, privacy decision, or release threshold | Request ID and input/prompt versions, model/version, token counts, latency and response status |
Pydantic |
Validate structured model output against typed fields and constraints | Semantic correctness, missing-evidence policy, or side-effect authorization | Validation result, rejected payload, repair or escalation path |
PyPDFLoader / document loader |
Extract page text and preserve document/page identity | Source reliability, parsing correctness, access policy, or deletion lifecycle | Document hash, parser version, page count, extraction warnings |
| Recursive text splitter | Create bounded overlapping chunks | The right semantic unit, overlap policy, or retrieval quality | Source document/page and chunk IDs, split positions, token counts and splitter configuration |
| BGE embedding model | Map queries and chunks into a retrieval vector space | Domain fit, access filtering, index freshness, or semantic truth | Model revision, vector dimension, normalization, benchmark slice |
FAISS |
Persist and search a vector index | Document authorization, metadata correctness, freshness, or relevance | Index version, build manifest, candidate IDs and distances |
| BM25 / lexical index | Retrieve exact names, identifiers, and rare terms | Synonym coverage or grounded answer generation | Analyzer configuration, term scores, corpus/index version |
| BGE reranker | Reorder a candidate set with query-document interaction | Recall lost before reranking or final answer faithfulness | Candidate set, reranker revision, before/after ranks, latency |
Ragas |
Compute configured RAG measures, including model-assisted metrics | Metric validity for the domain, judge calibration, or the final release decision | Pinned package version, judge model, metric inputs, raw result |
LangGraph |
Define graph state and routing, with optional configured checkpointing and resume | Business policy, tool permissions, state schema quality, or completion criteria | Graph version, node transitions, checkpoint ID, final state |
FastMCP / MCP server |
Expose typed tools, resources, and prompts through MCP messages between a host’s client and a server | Host consent, credential scope, server trust, or downstream authorization | Protocol/version, capability negotiation, principal, arguments, result provenance |
| Container or code sandbox | Constrain generated code execution and collect outputs | Whether code execution is necessary, safe inputs, or acceptable side effects | Image hash, resource limits, network/filesystem policy, exit status |
| Experiment tracker | Record dataset, prompt, model, code, metrics, and artifacts | The task definition, label quality, causal interpretation, or release decision | Immutable run ID, linked data, prompt, model, code and result versions, comparison and approval record |
| Telemetry and tracing | Observe latency, errors, token/tool use, and state transitions | The meaning of success, which private data may be logged, or retention policy | Trace ID, redaction status, sampled events, metric definitions |