55bca92691
Core features: - Class 1 T.30 protocol: full send/receive implementation - HDLC: DLE-stuffing, FCS strip, USR5637 bit-reversal handling - T.4 MH encoder/decoder (1728px A4 standard) - Document pipeline: PDF (Ghostscript), PNG, TIFF input - Width clamping: US Letter 1734px → 1728px fax standard - Cover page: CJK rasterization (TW/CN/JP/EN), TIFF + HTML output - OCR verification: Tesseract 5 with eng+chi_tra, CJK space-tolerant - API server (axum): health, send, jobs, cover, retry, cancel - Background worker: auto-poll queue, speed fallback, retry policy - Modem detection, pool management Real-world test results (2026-07-23): - V90 → 25153038: 4 pages, V.17 12000 bps, 2:33 ✅ - USR5637 → 25153038: 4 pages, V.17 12000 bps, 2:26 ✅ - Both faxes confirmed received on remote machine Tested: loopback (100% pixel match), multi-page, all input formats, cover pages, OCR verify, API endpoints, worker processing. 13 unit tests pass, 0 new clippy warnings.
597 lines
14 KiB
Markdown
597 lines
14 KiB
Markdown
# OCR 优化建议
|
||
|
||
## 🎯 优化目标
|
||
|
||
**当前问题:**
|
||
- English OCR: ✅ 99%准确度(完美)
|
||
- Chinese OCR: ⚠️ 60%准确度(需改进)
|
||
- Processing time: 可优化空间
|
||
|
||
**目标:**
|
||
- Chinese OCR准确度: 60% → **85%+**
|
||
- Processing speed: 保持或提升
|
||
- 商业部署: 生产就绪
|
||
|
||
---
|
||
|
||
## 📊 优化方案对比
|
||
|
||
### 方案 1: 安装更好的字体数据包(推荐)⭐⭐⭐⭐⭐
|
||
|
||
**优势:**
|
||
```
|
||
✅ 最简单
|
||
✅ 免费
|
||
✅ 快速提升准确度
|
||
✅ 无需代码修改
|
||
```
|
||
|
||
**实施:**
|
||
|
||
```bash
|
||
# 1. 下载最佳语言数据包
|
||
curl -L -o /opt/homebrew/share/tessdata/chi_tra_best.traineddata \
|
||
https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_tra.traineddata
|
||
|
||
curl -L -o /opt/homebrew/share/tessdata/chi_sim_best.traineddata \
|
||
https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_sim.traineddata
|
||
|
||
curl -L -o /opt/homebrew/share/tessdata/jpn_best.traineddata \
|
||
https://github.com/tesseract-ocr/tessdata_best/raw/main/jpn.traineddata
|
||
|
||
# 2. 使用最佳数据包
|
||
tesseract input.png stdout -l chi_tra_best
|
||
|
||
# 3. 验证改进
|
||
tesseract --list-langs
|
||
```
|
||
|
||
**效果预估:**
|
||
```
|
||
Chinese Traditional: 60% → 85%
|
||
Chinese Simplified: 60% → 85%
|
||
Japanese: 60% → 85%
|
||
```
|
||
|
||
**成本:**
|
||
```
|
||
下载大小: ~100MB per language
|
||
加载时间: +200ms
|
||
内存占用: +150MB
|
||
```
|
||
|
||
---
|
||
|
||
### 方案 2: 图像预处理(高性价比)⭐⭐⭐⭐
|
||
|
||
**优势:**
|
||
```
|
||
✅ 提升所有语言准确度
|
||
✅ Rust实现
|
||
✅ 无额外依赖
|
||
```
|
||
|
||
**实施:**
|
||
|
||
```rust
|
||
// src/ocr/preprocess.rs
|
||
use image::{ImageBuffer, Luma};
|
||
|
||
pub struct ImagePreprocessor {
|
||
target_dpi: u32,
|
||
}
|
||
|
||
impl ImagePreprocessor {
|
||
pub fn preprocess_for_ocr(image: &[u8]) -> Result<Vec<u8>> {
|
||
let img = image::load_from_memory(image)?;
|
||
|
||
// 1. 提高分辨率
|
||
let img = img.resize_exact(204 * 3, 196 * 3, image::imageops::FilterType::Lanczos3);
|
||
|
||
// 2. 灰度化
|
||
let img = img.grayscale();
|
||
|
||
// 3. 二值化
|
||
let img = self.binarize(&img);
|
||
|
||
// 4. 噪点去除
|
||
let img = self.remove_noise(&img);
|
||
|
||
// 5. 边缘增强
|
||
let img = self.enhance_edges(&img);
|
||
|
||
Ok(img.to_bytes())
|
||
}
|
||
|
||
fn binarize(img: &DynamicImage) -> DynamicImage {
|
||
// Adaptive thresholding
|
||
let threshold = 128;
|
||
img.brighten(20).contrast(1.2)
|
||
}
|
||
|
||
fn remove_noise(img: &DynamicImage) -> DynamicImage {
|
||
// Apply Gaussian blur for noise removal
|
||
img.blur(1.0)
|
||
}
|
||
|
||
fn enhance_edges(img: &DynamicImage) -> DynamicImage {
|
||
// Sharpen edges for better text recognition
|
||
img.sharpen(3.0)
|
||
}
|
||
}
|
||
```
|
||
|
||
**集成:**
|
||
```rust
|
||
// 在OCR处理前预处理
|
||
let preprocessed = ImagePreprocessor::preprocess_for_ocr(&raw_data)?;
|
||
let ocr_result = ocr_processor.process_data(&preprocessed, "tiff")?;
|
||
```
|
||
|
||
**效果预估:**
|
||
```
|
||
All languages: +10-15% accuracy
|
||
Processing: +50ms preprocessing
|
||
Quality: High
|
||
```
|
||
|
||
---
|
||
|
||
### 方案 3: 参数优化(低成本)⭐⭐⭐⭐⭐
|
||
|
||
**优势:**
|
||
```
|
||
✅ 免费
|
||
✅ 快速
|
||
✅ 无需额外安装
|
||
```
|
||
|
||
**实施:**
|
||
|
||
```rust
|
||
// src/ocr/mod.rs
|
||
impl OcrProcessor {
|
||
pub fn new_optimized() -> Self {
|
||
Self {
|
||
config: OcrConfig {
|
||
// 使用 LSTM 神经网络引擎(最准确)
|
||
oem: OcrEngineMode::NeuralNetLstmOnly,
|
||
|
||
// 自动页面分割(最灵活)
|
||
psm: PageSegMode::Auto,
|
||
|
||
// 高DPI传真图像
|
||
dpi: 204, // 或更高: 300
|
||
|
||
// 多语言组合
|
||
language: OcrLanguage::Multi(vec!["chi_tra", "eng"]),
|
||
},
|
||
}
|
||
}
|
||
|
||
pub fn process_with_optimization(&self, image_path: &Path) -> Result<OcrResult> {
|
||
let mut args = vec![
|
||
image_path.to_string_lossy().to_string(),
|
||
"stdout".to_string(),
|
||
|
||
// 最佳参数
|
||
"-l".to_string(), self.get_best_language_combo(),
|
||
"--dpi".to_string(), "204".to_string(),
|
||
"--oem".to_string(), "1".to_string(), // LSTM only
|
||
"--psm".to_string(), "3".to_string(), // Auto
|
||
|
||
// 高级参数
|
||
"--dpi".to_string(), "300".to_string(), // 提高DPI
|
||
"quiet".to_string(),
|
||
];
|
||
|
||
// 添加语言特定配置
|
||
args.push("-c".to_string());
|
||
args.push("tessedit_char_whitelist=0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz".to_string());
|
||
|
||
// 执行OCR
|
||
let output = Command::new(&self.tesseract_path)
|
||
.args(&args)
|
||
.output()?;
|
||
|
||
Ok(self.parse_result(output))
|
||
}
|
||
|
||
fn get_best_language_combo(&self) -> String {
|
||
match &self.config.language {
|
||
OcrLanguage::ChineseTraditional => "chi_tra_best+eng",
|
||
OcrLanguage::ChineseSimplified => "chi_sim_best+eng",
|
||
OcrLanguage::Japanese => "jpn_best+eng",
|
||
_ => "eng",
|
||
}
|
||
}
|
||
}
|
||
```
|
||
|
||
**最佳参数:**
|
||
```bash
|
||
# Tesseract参数优化
|
||
--oem 1 # LSTM神经网络(最准确)
|
||
--psm 3 # 自动页面分割(最灵活)
|
||
--dpi 300 # 高分辨率
|
||
-l chi_tra_best # 最佳语言包
|
||
```
|
||
|
||
**效果预估:**
|
||
```
|
||
Chinese: +5-10% accuracy
|
||
Free: Yes
|
||
Time: +50ms
|
||
```
|
||
|
||
---
|
||
|
||
### 方案 4: 字体优化(核心)⭐⭐⭐⭐⭐
|
||
|
||
**问题根源:**
|
||
```
|
||
中文OCR准确度低的原因:
|
||
1. 测试图片使用Helvetica字体(不支持中文)
|
||
2. PIL默认字体不支持中文字符
|
||
3. 需要使用专门的中文字体
|
||
```
|
||
|
||
**解决方案:**
|
||
|
||
```python
|
||
# scripts/cover_multilang.py
|
||
def create_chinese_cover_with_font(text, output_path):
|
||
from PIL import Image, ImageDraw, ImageFont
|
||
|
||
W, H = 1728, 2291
|
||
img = Image.new('L', (W, H), 255)
|
||
draw = ImageDraw.Draw(img)
|
||
|
||
# 使用中文字体(关键)
|
||
CHINESE_FONT_PATHS = [
|
||
'/System/Library/AssetsV2/com_apple_MobileAsset_Font8/86ba2c91f017a3749571a82f2c6d890ac7ffb2fb.asset/AssetData/PingFang.ttc',
|
||
'/System/Library/Fonts/STHeiti Medium.ttc',
|
||
'/usr/share/fonts/truetype/wqy/wqy-zenhei.ttc', # Linux
|
||
'/usr/share/fonts/opentype/noto/NotoSansCJK-Regular.ttc', # Linux
|
||
]
|
||
|
||
font_path = None
|
||
for path in CHINESE_FONT_PATHS:
|
||
if os.path.exists(path):
|
||
font_path = path
|
||
break
|
||
|
||
if font_path:
|
||
if 'PingFang' in font_path:
|
||
# PingFang字体索引(关键)
|
||
font_title = ImageFont.truetype(font_path, 72, index=10) # TC-Semibold
|
||
font_bold = ImageFont.truetype(font_path, 44, index=6) # TC-Medium
|
||
font_normal = ImageFont.truetype(font_path, 44, index=2) # TC-Light
|
||
else:
|
||
font_title = ImageFont.truetype(font_path, 72)
|
||
font_bold = ImageFont.truetype(font_path, 44)
|
||
font_normal = ImageFont.truetype(font_path, 44)
|
||
else:
|
||
# 下载并安装字体(可选)
|
||
print("Warning: No Chinese font found. Downloading...")
|
||
# 可以在这里添加自动下载逻辑
|
||
|
||
# 绘制中文文本
|
||
lines = text.split('\n')
|
||
y = 100
|
||
for line in lines:
|
||
draw.text((100, y), line, fill=0, font=font_normal)
|
||
y += 60
|
||
|
||
img.save(output_path, dpi=(204, 196))
|
||
```
|
||
|
||
**效果预估:**
|
||
```
|
||
Chinese OCR: 60% → 90%+ (字体匹配)
|
||
Key insight: OCR准确度取决于字体匹配度
|
||
```
|
||
|
||
---
|
||
|
||
### 方案 5: 多次OCR + 结果融合(高准确度)⭐⭐⭐
|
||
|
||
**原理:**
|
||
```
|
||
多次OCR处理不同参数,融合结果:
|
||
1. English only OCR
|
||
2. Chinese only OCR
|
||
3. Combined OCR
|
||
4. 选择最佳结果
|
||
```
|
||
|
||
**实施:**
|
||
|
||
```rust
|
||
pub fn multi_pass_ocr(&self, image_path: &Path) -> Result<OcrResult> {
|
||
let mut results = Vec::new();
|
||
|
||
// Pass 1: English only
|
||
let eng_result = self.process_with_language(image_path, "eng")?;
|
||
results.push(eng_result);
|
||
|
||
// Pass 2: Chinese Traditional
|
||
let chi_result = self.process_with_language(image_path, "chi_tra")?;
|
||
results.push(chi_result);
|
||
|
||
// Pass 3: Combined
|
||
let combined_result = self.process_with_language(image_path, "chi_tra+eng")?;
|
||
results.push(combined_result);
|
||
|
||
// 选择最佳结果(最高置信度)
|
||
let best_result = results.iter()
|
||
.max_by_key(|r| r.word_count)
|
||
.unwrap();
|
||
|
||
Ok(best_result.clone())
|
||
}
|
||
```
|
||
|
||
**效果预估:**
|
||
```
|
||
Accuracy: +10%
|
||
Processing: +300ms (3次处理)
|
||
Quality: High
|
||
```
|
||
|
||
---
|
||
|
||
### 方案 6: 机器学习后处理(高级)⭐⭐⭐
|
||
|
||
**原理:**
|
||
```
|
||
使用机器学习纠正OCR错误:
|
||
1. 收集常见错误样本
|
||
2. 训练纠错模型
|
||
3. 应用到OCR结果
|
||
```
|
||
|
||
**实施:**
|
||
|
||
```rust
|
||
// src/ocr/postprocess.rs
|
||
pub struct OcrPostProcessor {
|
||
// 常见错误纠正表
|
||
error_correction_map: HashMap<String, String>,
|
||
}
|
||
|
||
impl OcrPostProcessor {
|
||
pub fn correct_ocr_text(&self, text: &str) -> String {
|
||
let mut corrected = text.clone();
|
||
|
||
// 常见OCR错误纠正
|
||
let corrections = vec![
|
||
("O", "0"), // 数字0识别为字母O
|
||
("l", "1"), // 数字1识别为字母l
|
||
("S", "5"), // 数字5识别为字母S
|
||
("B", "8"), // 数字8识别为字母B
|
||
];
|
||
|
||
for (wrong, correct) in corrections {
|
||
// 根据上下文判断是否应该纠正
|
||
corrected = self.smart_replace(&corrected, wrong, correct);
|
||
}
|
||
|
||
corrected
|
||
}
|
||
|
||
fn smart_replace(&self, text: &str, wrong: &str, correct: &str) -> String {
|
||
// 智能替换:只在数字上下文中替换
|
||
// 例如:电话号码中的O→0,但单词中的O保留
|
||
let mut result = text.clone();
|
||
|
||
for (i, c) in text.char_indices() {
|
||
if c.to_string() == wrong {
|
||
// 检查周围字符是否是数字
|
||
let prev = text.chars().nth(i-1);
|
||
let next = text.chars().nth(i+1);
|
||
|
||
if prev.map(|p| p.is_digit(10)).unwrap_or(false) ||
|
||
next.map(|n| n.is_digit(10)).unwrap_or(false) {
|
||
result.replace_range(i..i+1, correct);
|
||
}
|
||
}
|
||
}
|
||
|
||
result
|
||
}
|
||
}
|
||
```
|
||
|
||
**效果预估:**
|
||
```
|
||
Accuracy: +5-8%
|
||
Processing: +20ms
|
||
Complexity: Medium
|
||
```
|
||
|
||
---
|
||
|
||
## 📋 实施优先级
|
||
|
||
### ⭐⭐⭐⭐⭐ 最高优先级(立即实施)
|
||
|
||
**1. 安装最佳语言数据包**
|
||
```bash
|
||
# 免费且效果最好
|
||
curl -L -o /opt/homebrew/share/tessdata/chi_tra_best.traineddata \
|
||
https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_tra.traineddata
|
||
```
|
||
|
||
**2. 使用正确的中文字体**
|
||
```python
|
||
# 确保封面生成使用PingFang字体
|
||
font_path = '/System/Library/.../PingFang.ttc'
|
||
font = ImageFont.truetype(font_path, 44, index=10)
|
||
```
|
||
|
||
**预期效果:**
|
||
```
|
||
Chinese OCR: 60% → 90%
|
||
Free: Yes
|
||
Time: 1-2 hours
|
||
```
|
||
|
||
---
|
||
|
||
### ⭐⭐⭐⭐ 高优先级(本周实施)
|
||
|
||
**3. 参数优化**
|
||
```rust
|
||
// 使用最佳Tesseract参数
|
||
--oem 1 --psm 3 --dpi 300
|
||
```
|
||
|
||
**4. 图像预处理**
|
||
```rust
|
||
// 实现图像预处理模块
|
||
ImagePreprocessor::preprocess_for_ocr(&raw_data)
|
||
```
|
||
|
||
**预期效果:**
|
||
```
|
||
All languages: +10-15%
|
||
Time: 2-3 days
|
||
```
|
||
|
||
---
|
||
|
||
### ⭐⭐⭐ 中优先级(下月实施)
|
||
|
||
**5. 多次OCR融合**
|
||
```rust
|
||
// 多次处理取最佳结果
|
||
multi_pass_ocr(&image_path)
|
||
```
|
||
|
||
**6. 后处理纠错**
|
||
```rust
|
||
// 智能纠错常见OCR错误
|
||
OcrPostProcessor::correct_ocr_text(&text)
|
||
```
|
||
|
||
---
|
||
|
||
## 🎯 预期最终效果
|
||
|
||
### 优化后准确度
|
||
|
||
| Language | Current | Optimized | Improvement |
|
||
|----------|---------|-----------|-------------|
|
||
| **English** | 99% | **99.5%** | +0.5% |
|
||
| **Chinese Traditional** | 60% | **90%** | **+30%** |
|
||
| **Chinese Simplified** | 60% | **90%** | **+30%** |
|
||
| **Japanese** | 60% | **85%** | **+25%** |
|
||
| **Multi-language** | 80% | **92%** | **+12%** |
|
||
|
||
### 优化后性能
|
||
|
||
| Metric | Current | Optimized | Impact |
|
||
|--------|---------|-----------|---------|
|
||
| **Processing** | 200ms | 250ms | +50ms |
|
||
| **Memory** | 50MB | 100MB | +50MB |
|
||
| **Quality** | Good | **Excellent** | ⭐⭐⭐⭐⭐ |
|
||
|
||
---
|
||
|
||
## 📝 实施步骤
|
||
|
||
### Step 1: 立即实施(1小时)
|
||
|
||
```bash
|
||
#!/bin/bash
|
||
# scripts/install_best_ocr.sh
|
||
|
||
echo "Installing best Tesseract language packs..."
|
||
|
||
# 下载最佳语言包
|
||
curl -L -o /opt/homebrew/share/tessdata/chi_tra_best.traineddata \
|
||
https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_tra.traineddata
|
||
|
||
curl -L -o /opt/homebrew/share/tessdata/chi_sim_best.traineddata \
|
||
https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_sim.traineddata
|
||
|
||
curl -L -o /opt/homebrew/share/tessdata/jpn_best.traineddata \
|
||
https://github.com/tesseract-ocr/tessdata_best/raw/main/jpn.traineddata
|
||
|
||
# 测试改进
|
||
tesseract --list-langs
|
||
|
||
echo "Testing improved OCR..."
|
||
cd /tmp
|
||
python3 /Users/accusys/telfax/scripts/create_test_image.py test_chinese_optimized.png chinese
|
||
tesseract test_chinese_optimized.png stdout -l chi_tra_best
|
||
|
||
echo "Optimization complete!"
|
||
```
|
||
|
||
### Step 2: 本周实施(2-3天)
|
||
|
||
```rust
|
||
// src/ocr/mod.rs - 添加优化配置
|
||
impl OcrProcessor {
|
||
pub fn with_best_config() -> Self {
|
||
Self {
|
||
config: OcrConfig {
|
||
language: OcrLanguage::Multi(vec!["chi_tra_best", "eng"]),
|
||
dpi: 300,
|
||
psm: PageSegMode::Auto,
|
||
oem: OcrEngineMode::NeuralNetLstmOnly,
|
||
},
|
||
}
|
||
}
|
||
}
|
||
|
||
// src/ocr/preprocess.rs - 添加预处理
|
||
pub fn preprocess_for_ocr(image: &[u8]) -> Result<Vec<u8>> {
|
||
// 图像预处理提升准确度
|
||
// ...
|
||
}
|
||
```
|
||
|
||
### Step 3: 下月实施(1-2周)
|
||
|
||
```rust
|
||
// src/ocr/postprocess.rs - 后处理纠错
|
||
pub fn correct_ocr_errors(text: &str) -> String {
|
||
// 智能纠错
|
||
// ...
|
||
}
|
||
```
|
||
|
||
---
|
||
|
||
## 💡 总结
|
||
|
||
**最优方案组合:**
|
||
|
||
```
|
||
方案1 + 方案2 + 方案4 = 最佳效果
|
||
- 最佳语言包(免费)
|
||
- 图像预处理(Rust)
|
||
- 正确字体(关键)
|
||
= Chinese OCR: 90%+
|
||
```
|
||
|
||
**实施建议:**
|
||
1. ⭐⭐⭐⭐⭐ 立即安装最佳语言包(1小时)
|
||
2. ⭐⭐⭐⭐⭐ 使用正确中文字体(关键)
|
||
3. ⭐⭐⭐⭐ 本周实现图像预处理(2-3天)
|
||
4. ⭐⭐⭐ 下月添加后处理纠错(1-2周)
|
||
|
||
**预期结果:**
|
||
```
|
||
✅ Chinese OCR: 90%+ 准确度
|
||
✅ All languages: 提升10-15%
|
||
✅ 生产就绪
|
||
✅ 商业部署合规
|
||
```
|
||
|
||
---
|
||
|
||
**优化建议完成!预期将中文OCR准确度提升至90%+** ✨ |