IronProw.IronHive 0.10.1

dotnet add package IronProw.IronHive --version 0.10.1
                    
NuGet\Install-Package IronProw.IronHive -Version 0.10.1
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="IronProw.IronHive" Version="0.10.1" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="IronProw.IronHive" Version="0.10.1" />
                    
Directory.Packages.props
<PackageReference Include="IronProw.IronHive" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add IronProw.IronHive --version 0.10.1
                    
#r "nuget: IronProw.IronHive, 0.10.1"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package IronProw.IronHive@0.10.1
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=IronProw.IronHive&version=0.10.1
                    
Install as a Cake Addin
#tool nuget:?package=IronProw.IronHive&version=0.10.1
                    
Install as a Cake Tool

iron-prow

safe-inference gateway — every LLM call's first gate.
provider selection (frontier ∨ LAN GpuStack) · guardrail · resilience — delivered as a standard Microsoft.Extensions.AI.IChatClient.

두 갈래 필수: frontier/LAN 게이트웨이(A) and local-provider safety(B).
역할·범위·설계 근거는 CHARTER.md 참조.
이미 자체 배선(provider·resilience·guardrail)을 가진 앱의 이관은 docs/ADOPTION.md 참조.

Packages

Package 역할
IronProw.Core 게이트웨이 계약, resilience, 선택 로직 — iyulab 구현 의존 0
IronProw.IronHive IronHive provider 어댑터 — OpenAI · Anthropic · GoogleAI (frontier) · GpuStack · OpenAI-Compatible/Ollama (LAN)
IronProw.FluxGuard FluxGuard guardrail 어댑터 (IGuard 구현)
IronProw.LMSupply 로컬 추론 안전 래퍼 (length-bounding + readiness gate)
dotnet add package IronProw.Core
dotnet add package IronProw.IronHive    # frontier provider adapters
dotnet add package IronProw.FluxGuard  # guardrail adapter
dotnet add package IronProw.LMSupply   # local-provider safety adapter

Quick Start

A. Frontier / LAN 게이트웨이

IronHive 기반 frontier provider + FluxGuard guardrail을 등록한다.
DI 컨테이너가 표준 IChatClient를 resolve하며, 게이트웨이가 selection·retry·입출력 검사를 처리한다(스트리밍 출력 검사의 시점은 아래 «What the guard inspects» 참조).

using IronProw.Core;
using IronProw.IronHive;
using IronProw.FluxGuard;
using Microsoft.Extensions.AI;
using Microsoft.Extensions.DependencyInjection;

var services = new ServiceCollection();   // Generic Host 라면 builder.Services
services.AddIronProw()
        .AddIronHiveOpenAI(
            id: "openai",
            priority: 10,
            modelId: "gpt-4o",
            configure: cfg => cfg.ApiKey = Environment.GetEnvironmentVariable("OPENAI_API_KEY")!)
        .AddIronHiveAnthropic(
            id: "anthropic",
            priority: 5,
            modelId: "claude-opus-4-5",
            configure: cfg => cfg.ApiKey = Environment.GetEnvironmentVariable("ANTHROPIC_API_KEY")!)
        .UseFluxGuard();  // Standard preset (L1 regex guards, offline)

// 표준 Microsoft.Extensions.AI.IChatClient 반환
await using var provider = services.BuildServiceProvider();
IChatClient chat = provider.GetRequiredService<IChatClient>();
try
{
    var response = await chat.GetResponseAsync("Hello");
    Console.WriteLine(response.Text);
}
catch (GuardException ex)   // 가드가 입력이나 출력을 막았다
{
    Console.WriteLine($"blocked: {ex.Reason}");
}

IronProw.IronHive는 다섯 가지 provider 어댑터를 제공한다:

  • AddIronHiveOpenAI · AddIronHiveAnthropic · AddIronHiveGoogleAI — frontier (ProviderKind.Frontier)
  • AddIronHiveGpuStack — LAN GpuStack (ProviderKind.Lan, key-optional). cfg => cfg.BaseUrl = "http://gpustack.lan:8080" 형태로 endpoint 지정.
  • AddIronHiveOpenAICompatible — LAN generic OpenAI-호환(Ollama·LMStudio·vLLM·llama.cpp server, ProviderKind.Lan, key-optional). 표준 /v1 API 표면을 노출하는 엔드포인트를 대상으로 하며 기본 endpoint는 Ollama의 http://localhost:11434. cfg => cfg.BaseUrl = "http://localhost:1234"(LMStudio)처럼 override.
services.AddIronProw()
        .AddIronHiveOpenAICompatible(
            id: "ollama",
            priority: 20,               // LAN 우선 — frontier보다 높게 두면 로컬 먼저 시도
            modelId: "llama3.2",
            configure: cfg => cfg.BaseUrl = "http://localhost:11434");  // key 불필요

UseFluxGuard() (파라미터 없음) 는 Standard preset(L1 regex, offline)을 적용한다.
UseFluxGuard(configure: b => …) 는 Standard preset 위에 FluxGuard 빌더 설정을 더한다(예: b.WithBlockThreshold(0.8)). 이미 만든 FluxGuard 인스턴스를 주입하려면 UseFluxGuard(IFluxGuard) 오버로드를 사용한다. 다른 가드는 UseGuard(sp => myGuard) 로 IGuard 를 직접 꽂는다.
fail mode: 기본은 fail-closed(불확실 verdict 차단). 가용성을 우선하는 소비자는 UseFluxGuard(failMode: FluxGuardFailMode.Open)으로 opt-in — Flagged/NeedsEscalation을 통과시킨다(정의된 Block은 mode 무관 항상 차단).

B. Local-provider safety (lm-supply / ONNX / DirectML)

로컬 추론을 iron-prow 안전 레이어로 감싼다. lm-supply 생성자(IGeneratorModel/ITextGenerator)를 그대로 넘기면 iron-prow가 내부 GeneratorChatClient 브리지로 IChatClient에 적응시킨다 — 소비자가 브리지를 직접 작성할 필요가 없다.

using IronProw.Core;
using IronProw.LMSupply;
using IronProw.FluxGuard;
using LMSupply.Generator;
using Microsoft.Extensions.AI;
using Microsoft.Extensions.DependencyInjection;

// generator: lm-supply 가 로드한 IGeneratorModel/ITextGenerator — 생명주기는 호출자 소유.
await using var generator = await LocalGenerator.LoadAsync("auto");

// probe: IReadinessProbe — 모델이 요청을 받을 수 있는지와 어떤 모델 id 가 있는지 보고한다.
//        LoadAsync 로 단일 모델을 직접 로드했다면 LazyReadinessProbe, GeneratorPool 을 쓰면 GeneratorPoolProbe.
var probe = new LazyReadinessProbe(() => true, [generator.GetModelInfo().ModelId]);

var services = new ServiceCollection();
services.AddIronProw()
        .AddLMSupplyLocal(
            id: "local-phi3",
            priority: 20,
            generator: generator,
            probe: probe,
            options: new LocalSafetyOptions { DefaultMaxOutputTokens = 1024 })
        .UseFluxGuard();

await using var provider = services.BuildServiceProvider();
IChatClient chat = provider.GetRequiredService<IChatClient>();

AddLMSupplyLocal이 제공하는 safety(호출마다 이 순서로 적용):

  • bridge — GeneratorChatClient가 lm-supply 생성자를 IChatClient로 적응(role 매핑, MaxOutputTokens→MaxNewTokens, sampler/tool 전파, streaming flatten). 이미 브리지된 IChatClient를 보유한 호출자(예: ironhive-host)는 AddLMSupplyLocal(..., IChatClient rawClient, ...) 오버로드를 쓸 수 있다.
  • readiness gate — IReadinessProbe.IsReadyAsync 가 false 면 InvalidOperationException. 게이트웨이는 이것을 그 provider 의 실패로 보고 다음 provider 로 넘어간다.
  • model-ID preflight — ChatOptions.ModelId 가 지정된 호출만 IReadinessProbe.GetAvailableModelIdsAsync 에 그 id 가 있는지 검증한다(없으면 InvalidOperationException). ⚠ 게이트웨이는 같은 ChatOptions 를 모든 provider 에 넘기므로, frontier 모델 id 를 지정한 호출이 local provider 로 fallback 되면 preflight 에서 실패하고 그 provider 의 건강 기록에도 실패로 남는다 — 혼합 게이트웨이에서는 ModelId 를 비워 두고 provider 별 기본 모델을 쓴다.
  • length-bounding — LocalSafetyOptions.DefaultMaxOutputTokens (미설정 호출에 자동 적용, 기본 512)
  • reasoning default — LocalSafetyOptions.DefaultReasoningEffort (미설정 호출에 자동 적용, 기본 ReasoningEffort.None). thinking 기본-on 모델(Gemma 4·Qwen3)은 작은 예산을 reasoning 에 전부 써 빈 답 + FinishReason.Length 를 돌려줄 수 있다 — 안전 래퍼의 계약은 «예산은 답에 쓴다»라 기본은 off. 모델 기본을 유지하려면 null.
  • reasoning 운반 — 브리지가 표준 ChatOptions.Reasoning 을 lm-supply ThinkingMode 로 번역한다(Effort.None→Off, 그 외→On, null→모델 기본). 모델이 낸 reasoning 은 TextReasoningContent 로 응답에 실린다(ReasoningOutput.None 이면 버림) — 빈 답이 왜 비었는지 소비자가 볼 수 있다.
  • 계측 — ChatResponse.ModelId·Usage(llama-server 의 prompt/completion 토큰; ONNX 는 null)·GetService<ChatClientMetadata>()(ProviderName = "LMSupply" — 경량 경로 BuildLocalSafeClient 가 돌려준 client 에서. 게이트웨이는 호출마다 provider 를 고르므로 provider 메타데이터를 내지 않는다(null); 어느 모델이 답했는지는 응답의 ModelId 로 본다).
경량 경로 — 단일 local provider (게이트웨이 없이)

폴백 대상 2번째 provider가 없는 local-first 단일 provider 소비자(예: textree)에게는 게이트웨이의 registry·selection·resilience 레이어가 전부 inert하다. 이 경우 BuildLocalSafeClient가 브리지+안전wrap만 조립한 plain IChatClient를 등록 없이 반환한다:

using IronProw.LMSupply;
using LMSupply.Generator;
using Microsoft.Extensions.AI;

await using var generator = await LocalGenerator.LoadAsync("auto");
var probe = new LazyReadinessProbe(() => true, [generator.GetModelInfo().ModelId]);

// 게이트웨이(AddIronProw/빌더/레지스트리) 없이 안전 래핑된 local client 직접 조립 — readiness·preflight·length-bound 만,
// 입출력 가드(FluxGuard)는 적용되지 않는다(가드가 필요하면 게이트웨이 경로를 쓴다).
IChatClient chat = LMSupplyExtensions.BuildLocalSafeClient(
    generator,                                                 // lm-supply ITextGenerator (호출자 소유)
    probe,                                                     // IReadinessProbe
    new LocalSafetyOptions { DefaultMaxOutputTokens = 1024 }); // 선택 (기본 512)

동일한 safety(readiness gate·preflight·length-bounding)를 받되 selection/fallback/resilience 오버헤드가 없다. 다중 provider·우선순위 선택·provider-level fallback이 필요해지면 AddLMSupplyLocal로 전환한다.

두 갈래 조합

두 갈래를 같은 DI 등록에서 조합할 수 있다. priority가 높은 provider가 먼저 선택되고, fallback 시 낮은 우선순위 provider로 강등된다.

services.AddIronProw()
        .AddIronHiveOpenAI("openai", priority: 10, "gpt-4o",
            cfg => cfg.ApiKey = Environment.GetEnvironmentVariable("OPENAI_API_KEY")!)
        .AddLMSupplyLocal("local", priority: 20, generator, probe)
        .UseFluxGuard()
        .Configure(opt =>
        {
            opt.EnableFallback = true;                              // 기본값
            opt.Resilience.MaxRetries = 3;                         // 기본값: 2
            opt.Resilience.BaseDelay = TimeSpan.FromMilliseconds(300); // 기본값: 200ms
            opt.Resilience.FailureThreshold = 3;                   // 연속 실패 N회면 cooldown (기본값: 3, 0 = 끔)
            opt.Resilience.Cooldown = TimeSpan.FromSeconds(30);    // cooldown 동안 후순위 (기본값: 30s)
            opt.Resilience.MaxRetryAfter = TimeSpan.FromSeconds(10); // provider 의 Retry-After 를 기다리는 상한 (기본값: 10s)
            opt.OnTransition = t =>                                 // retry/fallback/exhausted 이벤트 (UI 칩 등)
                Console.WriteLine($"[{t.Kind}] {t.ProviderId} ({t.ProviderIndex + 1}/{t.TotalProviders})");
        });

provider 건강 기억(0.4.0+): 게이트웨이는 provider 별 연속 실패를 기억한다. FailureThreshold 회 연속으로 강등 대상 실패(retry 소진·fallback-eligible)가 나면 그 provider 는 Cooldown 동안 후순위로 밀린다 — 제외가 아니라 후순위라 건강한 provider 가 하나도 없으면 여전히 시도되고, 한 번 성공하면 기록이 지워진다. 그 전엔 죽은 LAN provider 가 매 호출에 retry 예산(기본 2회 + backoff)을 물린 뒤에야 다음 provider 로 넘어갔다. 밀린 provider 는 OnTransition 에 ProwTransitionKind.Skipped 로 보고된다. EnableFallback = false 면 건강 기억도 라우팅에 쓰지 않는다(그 설정의 뜻이 «절대 provider 를 바꾸지 않는다»라서). 연결 거부·이름 해석 실패(HttpRequestError.ConnectionError/NameResolutionError)는 retry 가 아니라 즉시 fallback 으로 분류된다 — 엔드포인트가 바쁜 게 아니라 없는 것이라서.

HTTP 상태별 분류(0.5.0+): HTTP 실패는 예외 타입이 아니라 상태 코드로 분류된다 — OpenAI SDK 의 ClientResultException, HttpRequestException.StatusCode, ironhive 의 RateLimitException(AddIronHive* 가 등록하는 IronHiveHttpFailureReader) 모두 같은 규칙이다.

상태 분류
408 · 500 · 502 · 504 같은 provider 에서 retry
429 · 503 provider 가 retry 힌트(Retry-After / retry-after-ms)를 보냈으면 그만큼 기다려 retry(MaxRetryAfter 초과면 retry 없이 다음 provider), 없으면 즉시 다음 provider
그 밖(400 · 401 · 403 · 404 · 409 …) 다음 provider (한 provider 의 거절은 다른 provider 도 거절한다는 증거가 아니다)

retry 를 다 쓴 실패도 다음 provider 로 넘어간다. 상태에 도메인 의미를 주는 provider(예: 다운로드 승인 전까지 409 를 내는 로컬 provider)는 소비자가 IErrorClassifier 를 데코레이트해 그 코드만 다르게 분류한다. 다른 예외 타입이 HTTP 실패를 나르면 IHttpFailureReader 를 TryAddEnumerable 로 등록한다.

using IronProw.Core;
using Microsoft.Extensions.DependencyInjection;

// AddIronProw() 뒤에 등록하면 기본 분류기를 대신한다 — 409 만 다르게, 나머지는 기본 규칙에 맡긴다.
services.AddSingleton<IErrorClassifier>(sp =>
    new DownloadPendingClassifier(new DefaultErrorClassifier(sp.GetServices<IHttpFailureReader>())));

sealed class DownloadPendingClassifier(IErrorClassifier inner) : IErrorClassifier
{
    public ErrorClassification Classify(Exception exception)
        => exception is HttpRequestException { StatusCode: System.Net.HttpStatusCode.Conflict }
            ? ErrorClassification.FallbackEligible
            : inner.Classify(exception);
}

Provider 선택 — ProviderKind 와 IProviderSelector

기본 선택기(DefaultProviderSelector)는 priority 만 본다 — 높은 순서로 시도한다. 각 등록의 ProviderKind(Frontier·Lan·Local, 어댑터가 채운다)는 기본 선택기가 읽지 않는 메타데이터이고, 요청에 따라 순서를 바꾸고 싶은 소비자가 자기 IProviderSelector 에서 쓴다. AddIronProw() 뒤에 등록하면 기본 선택기를 대신한다. 선택기가 돌려준 순서에 건강 기억(cooldown 중인 provider 를 뒤로)이 적용된다.

using IronProw.Core;
using Microsoft.Extensions.AI;
using Microsoft.Extensions.DependencyInjection;

services.AddSingleton<IProviderSelector, ShortRequestsStayLocal>();

// 짧은 요청은 로컬 먼저, 나머지는 priority 순서.
sealed class ShortRequestsStayLocal : IProviderSelector
{
    public IReadOnlyList<ProviderRegistration> Order(ChatSelectionContext context)
    {
        var chars = context.Messages.Sum(m => m.Text.Length);
        return chars < 2_000
            ? [.. context.Candidates.OrderByDescending(c => c.Kind == ProviderKind.Local)]   // 안정 정렬 — 같은 Kind 안에선 priority 순서 유지
            : context.Candidates;
    }
}

OnTransition(선택)은 각 게이트웨이 전환(retry / fallback / exhausted / skipped)마다 호출되는 best-effort 콜백이다. 소비자가 어느 provider로 강등됐는지 UI에 표시(예: resilience 칩)할 수 있다. 콜백이 던지는 예외는 삼켜지며 추론을 절대 깨지 않는다. 미설정 시 동작은 기존과 동일(무보고).

스트리밍 동등성: retry·fallback은 GetResponseAsync와 GetStreamingResponseAsync 양쪽에 동일하게 적용된다. 스트리밍의 복원력 창은 "첫 ChatResponseUpdate가 yield되기 전" 이다 — 첫 청크 이전에 발생한 실패(예: OpenAI 호환 호출이 첫 MoveNextAsync에서 던지는 connection-refused / 404 / model-not-found)는 same-provider retry(Retryable) 또는 next-provider fallback(FallbackEligible)으로 처리된다. 첫 청크가 emit된 뒤의 실패는 provider를 바꾸면 이중 emit이 되므로 그대로 전파된다.

멀티테넌트 — per-tenant provider resolution

기본 AddProvider/AddLMSupplyLocal 경로는 provider 집합이 프로세스 수명 동안 고정인 소비자(데스크탑 에이전트, 단일 유저)를 위한 것이다. 워크스페이스마다 provider 집합·config·secret이 다른 멀티테넌트 서버 소비자는 AddTenantResolver로 per-tenant 게이트웨이를 런타임에 build한다.

using IronProw.Core;
using Microsoft.Extensions.AI;
using Microsoft.Extensions.DependencyInjection;

// startup — 단일 테넌트 AddProvider 경로와 병존(무회귀). 게이트웨이당 한 번만 호출한다(두 번째는 InvalidOperationException).
services.AddIronProw()
        .UseFluxGuard()
        .AddTenantResolver((sp, tenant) =>                           // 신규 표면
            sp.GetRequiredService<ProviderService>()                // consumer 구현
              .ResolveRegistrations(tenant));                       // per-workspace 집합 → ProviderRegistration[]

// per-request (요청 스코프에서 resolve)
IChatClient client = scopedSp.GetRequiredService<IIronProwFactory>().ForTenant(workspaceId);
await client.GetResponseAsync(messages, options, ct);                // guarded: select/retry/fallback/guard

// consumer 구현 — 워크스페이스의 provider 집합(요청 미들웨어가 미리 복호화해 둔 secret 을 sync 로 읽는다)
sealed class ProviderService
{
    public IReadOnlyList<ProviderRegistration> ResolveRegistrations(string workspaceId) => [];
}
  • ForTenant(tenant)은 해당 테넌트의 provider 집합으로 SelectingChatClient를 재조립한다 — selector/guard/classifier/options는 공유(재사용), registry만 per-tenant. tenant 키는 iron-prow에 opaque(resolver가 해석).
  • IIronProwFactory는 scoped로 등록되므로 요청 스코프에서 resolve해야 resolver·provider factory가 요청 범위 서비스(예: 복호화된 워크스페이스 secret)를 본다.
  • async는 상류에서: resolver와 ProviderRegistration.ClientFactory는 모두 sync다. 워크스페이스 secret의 async DB 로드·복호화는 consumer의 요청 미들웨어에서 수행해 scoped 서비스에 stash하고, resolver는 그것을 sync로 읽는다. (요청당 async 로드가 필수라면 향후 ForTenantAsync 오버로드가 순수 additive로 추가될 수 있다.)
  • 반환된 client는 매 요청 build(연결 없음·저비용)다. consumer가 provider factory에서 HttpClient 등 disposable을 쥐면 수명은 consumer 책임이다.
  • 단일 테넌트 경로(AddProvider/AddLMSupplyLocal, singleton IChatClient)는 완전 무변경으로 병존한다.

반복 퇴화 정지 (0.6.0+, opt-in)

작은 로컬 모델은 같은 단어·구절을 토큰 상한까지 반복하는 퇴화에 빠지곤 한다("concisely concisely concisely …"). WithDegenerationStop()은 어떤 IChatClient든(게이트웨이 포함) 감싸서, 출력 끝이 짧은 단위(≤ 60자, 글자 포함)의 4회 이상 연속 반복이 되면 생성을 멈춘다. 검사 대상은 출력의 마지막 240자다. Markdown 구조와 구두점 연속(----, | --- |, ====, ....)은 반복으로 보지 않는다. 판정은 좋은 답을 끊지 않는 쪽으로 치우쳐 있다.

using IronProw.Core;
using Microsoft.Extensions.AI;
using Microsoft.Extensions.DependencyInjection;

IChatClient chat = sp.GetRequiredService<IChatClient>().WithDegenerationStop(o => o.MinRepeats = 4);

// ChatClientBuilder 파이프라인에서는(ChatClientBuilder 는 Microsoft.Extensions.AI 패키지에 있다 — IronProw.Core 는 .Abstractions 만 참조):
IChatClient piped = new ChatClientBuilder(inner)
    .Use(c => new DegenerationStopChatClient(c, new DegenerationOptions { MinRepeats = 4 }))
    .Build();

DegenerationOptions — Window(검사할 출력 끝 길이, 기본 240자) · MaxPeriod(반복 단위 최대 길이, 기본 60자) · MinRepeats(연속 반복 횟수, 기본 4).

멈춤은 조용한 끝이 아니라 신호다. 스트리밍의 마지막 업데이트와 비스트리밍 응답의 FinishReason이 DegenerationStopChatClient.FinishReason("degeneration")이 되고, 반복된 단위는 AdditionalProperties[DegenerationStopChatClient.RepeatedUnitKey]에 실린다. 이미 내보낸 텍스트는 그대로 두고, 안쪽 스트림을 버려 생성을 취소한다. 비스트리밍 호출도 안쪽 스트림으로 받아 조기에 멈춘다. 판정만 필요하면 DegenerationDetector.FindRepeatingUnit(text)를 쓴다.

예방 쪽은 로컬 경로(IronProw.LMSupply)의 샘플러다. ChatOptions.FrequencyPenalty/PresencePenalty가 전달되고, lm-supply 고유의 repetition_penalty(기본 1.1)는 ChatOptions.AdditionalProperties["repetition_penalty"]로 준다.

Crash-fallback 제한

LocalSafetyChatClient(갈래 B)는 IReadinessProbe로 로컬 추론 불가를 감지하고, 게이트웨이 SelectingChatClient의 provider-level fallback으로 승격한다. 이것은 게이트웨이 수준 fallback(M2-4 범위)이다.

진짜 하드웨어 수준 fallback(예: ONNX GenAI DirectML → CPU)은 upstream lm-supply의 책임이다.
lm-supply 0.35.x 기준, ONNX GenAI 생성 경로는 DirectML → CPU 폴백이 없다(임베딩 경로·llama-server 경로는 보유). iron-prow는 이 gap을 감싸지 않는다 — upstream이 수정되면 iron-prow 변경 없이 이득이 흡수된다.

두 갈래 필수 설계

iron-prow는 두 시나리오를 독립적이면서도 조합 가능하게 커버한다:

갈래 진입점 책임
A. Frontier / LAN AddIronHiveOpenAI · AddIronHiveAnthropic · AddIronHiveGoogleAI · AddIronHiveGpuStack · AddIronHiveOpenAICompatible provider 레지스트리, 우선순위 선택, retry, provider-level fallback, 전환 이벤트(OnTransition)
B. local-provider safety AddLMSupplyLocal model-ID preflight, readiness gate, length-bounding, crash→fallback

어느 한 갈래만 구현하면 수요의 절반을 놓친다 (CHARTER §기능 표면).
UseFluxGuard()는 두 갈래 공통 — 관문에서 일괄 적용된다.

⚠️ The guard is opt-in. AddIronProw() installs a default NullGuard that allows all traffic. A gateway without UseFluxGuard() (or a custom UseGuard(...)) performs NO input/output guardrail checks. Always register a guard in production.

What the guard inspects, and when. A blocked input or output throws GuardException (Reason says why).

  • GetResponseAsync — the input before the call, the response after it; a blocked response is never returned.
  • GetStreamingResponseAsync — the input before the call, the aggregated output when the stream ends. Chunks arrive as they are generated, so a blocked output ends the stream with GuardException after its chunks were yielded: a streaming consumer discards or retracts what it showed when the exception arrives. (0.9.0+; before, streamed output was never inspected.)
  • A consumer that stops reading early — WithDegenerationStop() around the gateway, which also serves GetResponseAsync through the stream — still gets the output inspected, on disposal. (0.9.0+; before, WithDegenerationStop() around the gateway turned the output guard off for every call.)

See also

  • CHARTER.md — iron-prow 정체성, 범위, 의존 규칙, 로드맵 앵커

License

MIT

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
0.10.1 0 9/29/2026
0.10.0 0 9/29/2026
0.9.0 0 9/29/2026
0.8.14 0 9/29/2026
0.8.13 0 9/29/2026
0.8.12 0 9/28/2026
0.8.11 38 9/28/2026
0.8.10 39 9/28/2026
0.8.9 39 9/28/2026
0.8.8 39 9/28/2026
0.8.7 35 9/28/2026
0.8.6 41 9/27/2026
0.8.5 44 9/27/2026
0.8.4 42 9/27/2026
0.8.3 38 9/27/2026
0.8.2 39 9/27/2026
0.8.1 49 9/27/2026
0.8.0 47 9/27/2026
0.7.3 49 9/26/2026
0.7.2 48 9/26/2026
Loading failed