BlueBird.Pinyins 1.0.0

There is a newer version of this package available.
See the version list below for details.
dotnet add package BlueBird.Pinyins --version 1.0.0
                    
NuGet\Install-Package BlueBird.Pinyins -Version 1.0.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="BlueBird.Pinyins" Version="1.0.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="BlueBird.Pinyins" Version="1.0.0" />
                    
Directory.Packages.props
<PackageReference Include="BlueBird.Pinyins" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add BlueBird.Pinyins --version 1.0.0
                    
#r "nuget: BlueBird.Pinyins, 1.0.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package BlueBird.Pinyins@1.0.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=BlueBird.Pinyins&version=1.0.0
                    
Install as a Cake Addin
#tool nuget:?package=BlueBird.Pinyins&version=1.0.0
                    
Install as a Cake Tool

BlueBird.Pinyins

A zero-dependency .NET library for converting between Chinese characters and Pinyin. Supports Chinese-to-Pinyin, Pinyin-to-Chinese, and initial letter extraction.

Quick Start

using BlueBird.Pinyins;

// Chinese to Pinyin
Pinyin.GetPinyin("中国");         // "zhongguo"
Pinyin.GetPinyin("中国", " ");    // "zhong guo"

// Extract initial letters
Pinyin.GetInitials("你好");       // "nh"
Pinyin.GetInitials("你好", "-");  // "n-h"

// Pinyin to Chinese characters
Pinyin.GetChineseText("zhong");   // "中忠钟终盅..."

API Reference

Method Parameters Returns Description
GetPinyin(string?, string?) text, separator (optional) string? Returns pinyin string, null input returns null
GetInitials(string?, string?) text, separator (optional) string? Returns initial letters string, null input returns null
GetPinyin(char) character string Returns pinyin for a single character
GetChineseText(string?) pinyin string? Returns matching Chinese characters, null input returns null

Known Limitations

  • No multi-pronunciation (多音字) support.
  • No traditional Chinese (繁体字) support.
  • Characters not in the pinyin data table (letters, digits, punctuation, etc.) are returned as-is.

Architecture

Data Storage

All pinyin data is embedded as C# source code — zero external dependencies. The library uses a two-layer static index structure:

┌─────────────────────────────────────────┐
│           PyCode.Codes (399 entries)    │
│  ┌───────────────────────────────────┐  │
│  │ "a     :阿啊吖嗄腌锕"              │  │  Fixed format:
│  │ "ai    :爱埃碍矮挨唉..."           │  │  [6-char pinyin + 1-char sep + N chars]
│  │ "zhong :中忠钟终盅..."             │  │
│  │ "zuo   :作做左座坐昨..."           │  │  Separator: ":" for short pinyins, space for others
│  └───────────────────────────────────┘  │
└─────────────────────────────────────────┘

┌─────────────────────────────────────────┐
│         PyHash.Hashes (500 buckets)     │
│  ┌───────────────────────────────────┐  │
│  │ [0]  → {10, 31, 40, 65, ...}      │  │  Each bucket contains indices into PyCode
│  │ [1]  → {7, 125, 139, ...}         │  │
│  │ ...                               │  │  Hash function: (uint)ch % 500
│  │ [499]→ {36, 59, 84, ...}          │  │
│  └───────────────────────────────────┘  │
└─────────────────────────────────────────┘

Chinese → Pinyin

Input: '中' (Unicode 20013)
  ↓
Hash bucket: 20013 % 500 = 13
  ↓
Scan entries in bucket [13]:
  → Search for '中' in the character portion of each entry
  → Found → Extract first 6 chars "zhong ", TrimEnd → "zhong"
  → Not found → Return '中' as-is

Key design:

  • Hash bucketing: Characters are evenly distributed into 500 buckets by Unicode codepoint, avoiding full-table scans.
  • Non-Chinese fallback: Unmatched characters are returned as-is.

Pinyin → Chinese

Input: "zhong"
  ↓
Normalize: Trim() + ToLower()
  ↓
Sequential scan of PyCode.Codes (399 entries):
  → Match condition: StartsWith("zhong ") or StartsWith("zhong:")
  → Found → Extract character portion after position 7, return
  → Not found → Return empty string

Key design:

  • Fixed-width pinyin field: Pinyin is always padded to 6 characters (right-padded with spaces).

汉字与拼音转换工具库

快速开始

using BlueBird.Pinyins;

// 汉字转拼音
Pinyin.GetPinyin("中国");         // "zhongguo"
Pinyin.GetPinyin("中国", " ");    // "zhong guo"

// 提取拼音首字母
Pinyin.GetInitials("你好");       // "nh"
Pinyin.GetInitials("你好", "-");  // "n-h"

// 拼音转汉字
Pinyin.GetChineseText("zhong");   // "中忠钟终盅..."

API 说明

方法 参数 返回值 说明
GetPinyin(string?, string?) 文本, 分隔符(可选) string? 返回拼音串,null 输入返回 null
GetInitials(string?, string?) 文本, 分隔符(可选) string? 返回拼音首字母串,null 输入返回 null
GetPinyin(char) 字符 string 返回单个字符的拼音
GetChineseText(string?) 拼音 string? 返回对应汉字列表,null 输入返回 null

已知限制

  • 不支持多音字。
  • 不支持繁体字。
  • 不在拼音数据表中的字符(字母、数字、标点等)按原样返回。

技术原理

数据存储结构

本库采用 双层静态索引表 结构,零外部依赖,所有数据内嵌为 C# 源代码。

┌─────────────────────────────────────────┐
│           PyCode.Codes (399 条)         │
│  ┌───────────────────────────────────┐  │
│  │ "a     :阿啊吖嗄腌锕"              │  │  每条固定格式:
│  │ "ai    :爱埃碍矮挨唉..."           │  │  [6位拼音 + 1位分隔符 + N个汉字]
│  │ "zhong :中忠钟终盅..."             │  │
│  │ "zuo   :作做左座坐昨..."           │  │  分隔符: 为短拼音用":",其余用空格
│  └───────────────────────────────────┘  │
└─────────────────────────────────────────┘

┌─────────────────────────────────────────┐
│         PyHash.Hashes (500 个桶)        │
│  ┌───────────────────────────────────┐  │
│  │ [0]  → {10, 31, 40, 65, ...}      │  │  每个桶包含指向 PyCode 的索引
│  │ [1]  → {7, 125, 139, ...}         │  │
│  │ ...                               │  │  哈希函数: (uint)ch % 500
│  │ [499]→ {36, 59, 84, ...}          │  │
│  └───────────────────────────────────┘  │
└─────────────────────────────────────────┘
汉字转拼音
输入: '中' (Unicode 20013)
  ↓
计算哈希桶: 20013 % 500 = 13
  ↓
扫描桶 [13] 中每个 PyCode 索引
  → 在索引对应条目的汉字部分查找 '中'
  → 找到 → 提取前6位 "zhong ",TrimEnd → "zhong"
  → 未找到 → 遍历完桶后原样返回 '中'

关键设计:

  • 哈希分桶:用字符的 Unicode 码点对 500 取模,将常见汉字均匀分布到 500 个桶中,避免全表扫描
  • 非汉字回退:未命中任何条目时,直接返回字符本身(适用于字母、数字、标点等)
拼音转汉字
输入: "zhong"
  ↓
标准化: Trim() + ToLower()
  ↓
顺序扫描 PyCode.Codes (399条):
  → 匹配条件: StartsWith("zhong ") 或 StartsWith("zhong:")
  → 找到 → 提取第7位之后的汉字部分,返回
  → 未找到 → 返回空字符串

关键设计:

  • 定长分隔符:拼音字段固定6字符(不足右侧补空格)
Product Compatible and additional computed target framework versions.
.NET net8.0 is compatible.  net8.0-android was computed.  net8.0-browser was computed.  net8.0-ios was computed.  net8.0-maccatalyst was computed.  net8.0-macos was computed.  net8.0-tvos was computed.  net8.0-windows was computed.  net9.0 was computed.  net9.0-android was computed.  net9.0-browser was computed.  net9.0-ios was computed.  net9.0-maccatalyst was computed.  net9.0-macos was computed.  net9.0-tvos was computed.  net9.0-windows was computed.  net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.
  • net10.0

    • No dependencies.
  • net8.0

    • No dependencies.

NuGet packages

This package is not used by any NuGet packages.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
1.3.0 110 8/30/2026
1.2.0 166 6/7/2026
1.1.0 220 5/30/2026
1.0.0 145 5/24/2026

Chinese character ↔ Pinyin conversion: GetPinyin, GetInitials, GetChineseText.