Files
telfax/docs/OCR_OPTIMIZATION_GUIDE.md
Warren 55bca92691 V1.0: Class 1 fax — real-world 4-page send to external number confirmed
Core features:
- Class 1 T.30 protocol: full send/receive implementation
- HDLC: DLE-stuffing, FCS strip, USR5637 bit-reversal handling
- T.4 MH encoder/decoder (1728px A4 standard)
- Document pipeline: PDF (Ghostscript), PNG, TIFF input
- Width clamping: US Letter 1734px → 1728px fax standard
- Cover page: CJK rasterization (TW/CN/JP/EN), TIFF + HTML output
- OCR verification: Tesseract 5 with eng+chi_tra, CJK space-tolerant
- API server (axum): health, send, jobs, cover, retry, cancel
- Background worker: auto-poll queue, speed fallback, retry policy
- Modem detection, pool management

Real-world test results (2026-07-23):
- V90 → 25153038: 4 pages, V.17 12000 bps, 2:33 ✅
- USR5637 → 25153038: 4 pages, V.17 12000 bps, 2:26 ✅
- Both faxes confirmed received on remote machine

Tested: loopback (100% pixel match), multi-page, all input formats,
cover pages, OCR verify, API endpoints, worker processing.
13 unit tests pass, 0 new clippy warnings.
2026-07-24 18:47:15 +08:00

597 lines
14 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# OCR 优化建议
## 🎯 优化目标
**当前问题:**
- English OCR: ✅ 99%准确度(完美)
- Chinese OCR: ⚠️ 60%准确度(需改进)
- Processing time: 可优化空间
**目标:**
- Chinese OCR准确度: 60% → **85%+**
- Processing speed: 保持或提升
- 商业部署: 生产就绪
---
## 📊 优化方案对比
### 方案 1: 安装更好的字体数据包(推荐)⭐⭐⭐⭐⭐
**优势:**
```
✅ 最简单
✅ 免费
✅ 快速提升准确度
✅ 无需代码修改
```
**实施:**
```bash
# 1. 下载最佳语言数据包
curl -L -o /opt/homebrew/share/tessdata/chi_tra_best.traineddata \
https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_tra.traineddata
curl -L -o /opt/homebrew/share/tessdata/chi_sim_best.traineddata \
https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_sim.traineddata
curl -L -o /opt/homebrew/share/tessdata/jpn_best.traineddata \
https://github.com/tesseract-ocr/tessdata_best/raw/main/jpn.traineddata
# 2. 使用最佳数据包
tesseract input.png stdout -l chi_tra_best
# 3. 验证改进
tesseract --list-langs
```
**效果预估:**
```
Chinese Traditional: 60% → 85%
Chinese Simplified: 60% → 85%
Japanese: 60% → 85%
```
**成本:**
```
下载大小: ~100MB per language
加载时间: +200ms
内存占用: +150MB
```
---
### 方案 2: 图像预处理(高性价比)⭐⭐⭐⭐
**优势:**
```
✅ 提升所有语言准确度
✅ Rust实现
✅ 无额外依赖
```
**实施:**
```rust
// src/ocr/preprocess.rs
use image::{ImageBuffer, Luma};
pub struct ImagePreprocessor {
target_dpi: u32,
}
impl ImagePreprocessor {
pub fn preprocess_for_ocr(image: &[u8]) -> Result<Vec<u8>> {
let img = image::load_from_memory(image)?;
// 1. 提高分辨率
let img = img.resize_exact(204 * 3, 196 * 3, image::imageops::FilterType::Lanczos3);
// 2. 灰度化
let img = img.grayscale();
// 3. 二值化
let img = self.binarize(&img);
// 4. 噪点去除
let img = self.remove_noise(&img);
// 5. 边缘增强
let img = self.enhance_edges(&img);
Ok(img.to_bytes())
}
fn binarize(img: &DynamicImage) -> DynamicImage {
// Adaptive thresholding
let threshold = 128;
img.brighten(20).contrast(1.2)
}
fn remove_noise(img: &DynamicImage) -> DynamicImage {
// Apply Gaussian blur for noise removal
img.blur(1.0)
}
fn enhance_edges(img: &DynamicImage) -> DynamicImage {
// Sharpen edges for better text recognition
img.sharpen(3.0)
}
}
```
**集成:**
```rust
// 在OCR处理前预处理
let preprocessed = ImagePreprocessor::preprocess_for_ocr(&raw_data)?;
let ocr_result = ocr_processor.process_data(&preprocessed, "tiff")?;
```
**效果预估:**
```
All languages: +10-15% accuracy
Processing: +50ms preprocessing
Quality: High
```
---
### 方案 3: 参数优化(低成本)⭐⭐⭐⭐⭐
**优势:**
```
✅ 免费
✅ 快速
✅ 无需额外安装
```
**实施:**
```rust
// src/ocr/mod.rs
impl OcrProcessor {
pub fn new_optimized() -> Self {
Self {
config: OcrConfig {
// 使用 LSTM 神经网络引擎(最准确)
oem: OcrEngineMode::NeuralNetLstmOnly,
// 自动页面分割(最灵活)
psm: PageSegMode::Auto,
// 高DPI传真图像
dpi: 204, // 或更高: 300
// 多语言组合
language: OcrLanguage::Multi(vec!["chi_tra", "eng"]),
},
}
}
pub fn process_with_optimization(&self, image_path: &Path) -> Result<OcrResult> {
let mut args = vec![
image_path.to_string_lossy().to_string(),
"stdout".to_string(),
// 最佳参数
"-l".to_string(), self.get_best_language_combo(),
"--dpi".to_string(), "204".to_string(),
"--oem".to_string(), "1".to_string(), // LSTM only
"--psm".to_string(), "3".to_string(), // Auto
// 高级参数
"--dpi".to_string(), "300".to_string(), // 提高DPI
"quiet".to_string(),
];
// 添加语言特定配置
args.push("-c".to_string());
args.push("tessedit_char_whitelist=0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz".to_string());
// 执行OCR
let output = Command::new(&self.tesseract_path)
.args(&args)
.output()?;
Ok(self.parse_result(output))
}
fn get_best_language_combo(&self) -> String {
match &self.config.language {
OcrLanguage::ChineseTraditional => "chi_tra_best+eng",
OcrLanguage::ChineseSimplified => "chi_sim_best+eng",
OcrLanguage::Japanese => "jpn_best+eng",
_ => "eng",
}
}
}
```
**最佳参数:**
```bash
# Tesseract参数优化
--oem 1 # LSTM神经网络(最准确)
--psm 3 # 自动页面分割(最灵活)
--dpi 300 # 高分辨率
-l chi_tra_best # 最佳语言包
```
**效果预估:**
```
Chinese: +5-10% accuracy
Free: Yes
Time: +50ms
```
---
### 方案 4: 字体优化(核心)⭐⭐⭐⭐⭐
**问题根源:**
```
中文OCR准确度低的原因:
1. 测试图片使用Helvetica字体(不支持中文)
2. PIL默认字体不支持中文字符
3. 需要使用专门的中文字体
```
**解决方案:**
```python
# scripts/cover_multilang.py
def create_chinese_cover_with_font(text, output_path):
from PIL import Image, ImageDraw, ImageFont
W, H = 1728, 2291
img = Image.new('L', (W, H), 255)
draw = ImageDraw.Draw(img)
# 使用中文字体(关键)
CHINESE_FONT_PATHS = [
'/System/Library/AssetsV2/com_apple_MobileAsset_Font8/86ba2c91f017a3749571a82f2c6d890ac7ffb2fb.asset/AssetData/PingFang.ttc',
'/System/Library/Fonts/STHeiti Medium.ttc',
'/usr/share/fonts/truetype/wqy/wqy-zenhei.ttc', # Linux
'/usr/share/fonts/opentype/noto/NotoSansCJK-Regular.ttc', # Linux
]
font_path = None
for path in CHINESE_FONT_PATHS:
if os.path.exists(path):
font_path = path
break
if font_path:
if 'PingFang' in font_path:
# PingFang字体索引(关键)
font_title = ImageFont.truetype(font_path, 72, index=10) # TC-Semibold
font_bold = ImageFont.truetype(font_path, 44, index=6) # TC-Medium
font_normal = ImageFont.truetype(font_path, 44, index=2) # TC-Light
else:
font_title = ImageFont.truetype(font_path, 72)
font_bold = ImageFont.truetype(font_path, 44)
font_normal = ImageFont.truetype(font_path, 44)
else:
# 下载并安装字体(可选)
print("Warning: No Chinese font found. Downloading...")
# 可以在这里添加自动下载逻辑
# 绘制中文文本
lines = text.split('\n')
y = 100
for line in lines:
draw.text((100, y), line, fill=0, font=font_normal)
y += 60
img.save(output_path, dpi=(204, 196))
```
**效果预估:**
```
Chinese OCR: 60% → 90%+ (字体匹配)
Key insight: OCR准确度取决于字体匹配度
```
---
### 方案 5: 多次OCR + 结果融合(高准确度)⭐⭐⭐
**原理:**
```
多次OCR处理不同参数,融合结果:
1. English only OCR
2. Chinese only OCR
3. Combined OCR
4. 选择最佳结果
```
**实施:**
```rust
pub fn multi_pass_ocr(&self, image_path: &Path) -> Result<OcrResult> {
let mut results = Vec::new();
// Pass 1: English only
let eng_result = self.process_with_language(image_path, "eng")?;
results.push(eng_result);
// Pass 2: Chinese Traditional
let chi_result = self.process_with_language(image_path, "chi_tra")?;
results.push(chi_result);
// Pass 3: Combined
let combined_result = self.process_with_language(image_path, "chi_tra+eng")?;
results.push(combined_result);
// 选择最佳结果(最高置信度)
let best_result = results.iter()
.max_by_key(|r| r.word_count)
.unwrap();
Ok(best_result.clone())
}
```
**效果预估:**
```
Accuracy: +10%
Processing: +300ms (3次处理)
Quality: High
```
---
### 方案 6: 机器学习后处理(高级)⭐⭐⭐
**原理:**
```
使用机器学习纠正OCR错误:
1. 收集常见错误样本
2. 训练纠错模型
3. 应用到OCR结果
```
**实施:**
```rust
// src/ocr/postprocess.rs
pub struct OcrPostProcessor {
// 常见错误纠正表
error_correction_map: HashMap<String, String>,
}
impl OcrPostProcessor {
pub fn correct_ocr_text(&self, text: &str) -> String {
let mut corrected = text.clone();
// 常见OCR错误纠正
let corrections = vec![
("O", "0"), // 数字0识别为字母O
("l", "1"), // 数字1识别为字母l
("S", "5"), // 数字5识别为字母S
("B", "8"), // 数字8识别为字母B
];
for (wrong, correct) in corrections {
// 根据上下文判断是否应该纠正
corrected = self.smart_replace(&corrected, wrong, correct);
}
corrected
}
fn smart_replace(&self, text: &str, wrong: &str, correct: &str) -> String {
// 智能替换:只在数字上下文中替换
// 例如:电话号码中的O→0,但单词中的O保留
let mut result = text.clone();
for (i, c) in text.char_indices() {
if c.to_string() == wrong {
// 检查周围字符是否是数字
let prev = text.chars().nth(i-1);
let next = text.chars().nth(i+1);
if prev.map(|p| p.is_digit(10)).unwrap_or(false) ||
next.map(|n| n.is_digit(10)).unwrap_or(false) {
result.replace_range(i..i+1, correct);
}
}
}
result
}
}
```
**效果预估:**
```
Accuracy: +5-8%
Processing: +20ms
Complexity: Medium
```
---
## 📋 实施优先级
### ⭐⭐⭐⭐⭐ 最高优先级(立即实施)
**1. 安装最佳语言数据包**
```bash
# 免费且效果最好
curl -L -o /opt/homebrew/share/tessdata/chi_tra_best.traineddata \
https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_tra.traineddata
```
**2. 使用正确的中文字体**
```python
# 确保封面生成使用PingFang字体
font_path = '/System/Library/.../PingFang.ttc'
font = ImageFont.truetype(font_path, 44, index=10)
```
**预期效果:**
```
Chinese OCR: 60% → 90%
Free: Yes
Time: 1-2 hours
```
---
### ⭐⭐⭐⭐ 高优先级(本周实施)
**3. 参数优化**
```rust
// 使用最佳Tesseract参数
--oem 1 --psm 3 --dpi 300
```
**4. 图像预处理**
```rust
// 实现图像预处理模块
ImagePreprocessor::preprocess_for_ocr(&raw_data)
```
**预期效果:**
```
All languages: +10-15%
Time: 2-3 days
```
---
### ⭐⭐⭐ 中优先级(下月实施)
**5. 多次OCR融合**
```rust
// 多次处理取最佳结果
multi_pass_ocr(&image_path)
```
**6. 后处理纠错**
```rust
// 智能纠错常见OCR错误
OcrPostProcessor::correct_ocr_text(&text)
```
---
## 🎯 预期最终效果
### 优化后准确度
| Language | Current | Optimized | Improvement |
|----------|---------|-----------|-------------|
| **English** | 99% | **99.5%** | +0.5% |
| **Chinese Traditional** | 60% | **90%** | **+30%** |
| **Chinese Simplified** | 60% | **90%** | **+30%** |
| **Japanese** | 60% | **85%** | **+25%** |
| **Multi-language** | 80% | **92%** | **+12%** |
### 优化后性能
| Metric | Current | Optimized | Impact |
|--------|---------|-----------|---------|
| **Processing** | 200ms | 250ms | +50ms |
| **Memory** | 50MB | 100MB | +50MB |
| **Quality** | Good | **Excellent** | ⭐⭐⭐⭐⭐ |
---
## 📝 实施步骤
### Step 1: 立即实施(1小时)
```bash
#!/bin/bash
# scripts/install_best_ocr.sh
echo "Installing best Tesseract language packs..."
# 下载最佳语言包
curl -L -o /opt/homebrew/share/tessdata/chi_tra_best.traineddata \
https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_tra.traineddata
curl -L -o /opt/homebrew/share/tessdata/chi_sim_best.traineddata \
https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_sim.traineddata
curl -L -o /opt/homebrew/share/tessdata/jpn_best.traineddata \
https://github.com/tesseract-ocr/tessdata_best/raw/main/jpn.traineddata
# 测试改进
tesseract --list-langs
echo "Testing improved OCR..."
cd /tmp
python3 /Users/accusys/telfax/scripts/create_test_image.py test_chinese_optimized.png chinese
tesseract test_chinese_optimized.png stdout -l chi_tra_best
echo "Optimization complete!"
```
### Step 2: 本周实施(2-3天)
```rust
// src/ocr/mod.rs - 添加优化配置
impl OcrProcessor {
pub fn with_best_config() -> Self {
Self {
config: OcrConfig {
language: OcrLanguage::Multi(vec!["chi_tra_best", "eng"]),
dpi: 300,
psm: PageSegMode::Auto,
oem: OcrEngineMode::NeuralNetLstmOnly,
},
}
}
}
// src/ocr/preprocess.rs - 添加预处理
pub fn preprocess_for_ocr(image: &[u8]) -> Result<Vec<u8>> {
// 图像预处理提升准确度
// ...
}
```
### Step 3: 下月实施(1-2周)
```rust
// src/ocr/postprocess.rs - 后处理纠错
pub fn correct_ocr_errors(text: &str) -> String {
// 智能纠错
// ...
}
```
---
## 💡 总结
**最优方案组合:**
```
方案1 + 方案2 + 方案4 = 最佳效果
- 最佳语言包(免费)
- 图像预处理(Rust)
- 正确字体(关键)
= Chinese OCR: 90%+
```
**实施建议:**
1. ⭐⭐⭐⭐⭐ 立即安装最佳语言包(1小时)
2. ⭐⭐⭐⭐⭐ 使用正确中文字体(关键)
3. ⭐⭐⭐⭐ 本周实现图像预处理(2-3天)
4. ⭐⭐⭐ 下月添加后处理纠错(1-2周)
**预期结果:**
```
✅ Chinese OCR: 90%+ 准确度
✅ All languages: 提升10-15%
✅ 生产就绪
✅ 商业部署合规
```
---
**优化建议完成!预期将中文OCR准确度提升至90%+** ✨