RAGToolkit.AI.DataRetrieval 1.0.0

There is a newer version of this package available.
See the version list below for details.
dotnet add package RAGToolkit.AI.DataRetrieval --version 1.0.0
                    
NuGet\Install-Package RAGToolkit.AI.DataRetrieval -Version 1.0.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="RAGToolkit.AI.DataRetrieval" Version="1.0.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="RAGToolkit.AI.DataRetrieval" Version="1.0.0" />
                    
Directory.Packages.props
<PackageReference Include="RAGToolkit.AI.DataRetrieval" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add RAGToolkit.AI.DataRetrieval --version 1.0.0
                    
#r "nuget: RAGToolkit.AI.DataRetrieval, 1.0.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package RAGToolkit.AI.DataRetrieval@1.0.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=RAGToolkit.AI.DataRetrieval&version=1.0.0
                    
Install as a Cake Addin
#tool nuget:?package=RAGToolkit.AI.DataRetrieval&version=1.0.0
                    
Install as a Cake Tool

RAGToolkit.AI.DataRetrieval

管线的默认实现与内置处理器。

核心类型

类型 说明
RetrievalPipeline 管线:持有 QueryProcessors / ResultProcessors 两个列表,ProcessAsync 把它们串起来
VectorStoreRetriever 直接面向 MEVD VectorStoreCollection<TKey, TRecord> 的检索实现
RetrievalPipelineExtensions AsRetriever(collection, contentSelector) —— 把管线绑成 IRetriever 供 DI 使用
ReciprocalRankFusion RRF 融合(internal,多 variant 时由管线调用)
ChunkIdentity 去重用的块身份(internal)

内置处理器

Query Processors(检索前)

处理器 作用 依赖
NormalizationQueryProcessor Unicode NFKC 归一化(全角折半角)+ 空白折叠 + 可选大小写折叠 无
SynonymExpansionQueryProcessor 按术语表替换,每命中一次产出一个 variant 无
HydeQueryProcessor 让模型写一段"假答案",用它去检索稠密侧 一个生成模型
MultiQueryQueryProcessor 让模型给几种不同说法,各自检索后 RRF 融合 一个生成模型
RewritingQueryProcessor 让模型改写查询(稠密侧和关键词侧都变) 一个生成模型

Result Processors(检索后)

处理器 作用 元数据标志
ScoreThresholdResultProcessor 按分数阈值过滤 score_filtered
TopNResultProcessor 截断到前 N 条 trimmed
RerankingResultProcessor cross-encoder 重排 reranked
ParentContextResultProcessor 用相邻块扩展命中块 parent_context_expanded
LongContextReorderResultProcessor 最相关的放头尾、次要的放中间 reordered

每个处理器都会往 RetrievalResults.Metadata 写一个"本次是否起了作用"的布尔,链路里哪一环是空转的一眼可见。

处理器顺序

三条顺序是有意义的,写错了会静默失效:

1. HydeQueryProcessor 要放在任何 variant 扩展器之前。

它写的是 RetrievalProviderOptions.EmbeddingQueryText —— 那是每个查询一份的覆写,不是每 variant 一份。 查询已经被扩成多个 variant 时它不会动手,Metadata["hyde_skipped"] 会报 "several_variants"。

2. LongContextReorderResultProcessor 要放在最后。

它的输出故意不再按分数排序(最相关的在头尾)。后面再挂 TopNResultProcessor 会把顺序排回去,等于白做。

3. 融合之后不能再套相似度阈值。

多 variant 走 RRF 后,RetrievalChunk.Score 变成 rank 尺度 —— 顶部只有 1 / 61 ≈ 0.016,不再是相似度。 此时 ScoreThresholdResultProcessor(0.5) 会把结果全部清空。 Metadata["fused"] 报是否融合过,fusion_k 报 k 值。

接 LLM

三个 LLM 处理器都不挑模型,任意 IChatClient 实现(OpenAI / Azure / Ollama / DeepSeek / 通义 / 智谱 …)都能接:

using Microsoft.Extensions.AI;

IChatClient chat = /* 你的 chat client */;

// HyDE:让模型先写一段"假答案",再用它去检索稠密侧。
// shouldApply 可选 —— 短查询才值得扩写;长查询本来就具体,扩写反而会把
// 向量拉向模型的通用措辞,离语料更远。
pipeline.QueryProcessors.Add(new HydeQueryProcessor(
    async (question, cancellationToken) =>
        (await chat.GetResponseAsync(
            $"写一段能回答下面这个问题的文字,直接写,不要解释:{question}",
            cancellationToken: cancellationToken)).Text,
    shouldApply: query => query.Text.Length < 20));

// Multi-Query:让模型给几种不同说法,各自检索后用 RRF 融合。
pipeline.QueryProcessors.Add(new MultiQueryQueryProcessor(
    async (question, cancellationToken) =>
    {
        var response = await chat.GetResponseAsync(
            $"用 3 种不同的说法改写这个问题,每行一条,不要编号:{question}",
            cancellationToken: cancellationToken);

        return (response.Text ?? string.Empty)
            .Split('\n', StringSplitOptions.RemoveEmptyEntries | StringSplitOptions.TrimEntries);
    }));

零依赖的两个不需要模型:

pipeline.QueryProcessors.Add(new NormalizationQueryProcessor(foldCase: true));

pipeline.QueryProcessors.Add(new SynonymExpansionQueryProcessor(
    new Dictionary<string, IReadOnlyList<string>>
    {
        ["AI"] = ["Artificial Intelligence"],
        ["款号"] = ["商品编号"],
    }));

依赖

  • RAGToolkit.AI.DataRetrieval.Abstractions
  • Microsoft.Extensions.VectorData.Abstractions
  • Microsoft.Extensions.Logging.Abstractions

许可证

MIT —— 见 LICENSE。

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
1.1.1 0 10/11/2026
1.1.0 0 10/11/2026
1.0.0 42 10/9/2026