Files
telfax/docs/OCR_OPTIMIZATION_GUIDE.md
T
Warren 55bca92691 V1.0: Class 1 fax — real-world 4-page send to external number confirmed
Core features:
- Class 1 T.30 protocol: full send/receive implementation
- HDLC: DLE-stuffing, FCS strip, USR5637 bit-reversal handling
- T.4 MH encoder/decoder (1728px A4 standard)
- Document pipeline: PDF (Ghostscript), PNG, TIFF input
- Width clamping: US Letter 1734px → 1728px fax standard
- Cover page: CJK rasterization (TW/CN/JP/EN), TIFF + HTML output
- OCR verification: Tesseract 5 with eng+chi_tra, CJK space-tolerant
- API server (axum): health, send, jobs, cover, retry, cancel
- Background worker: auto-poll queue, speed fallback, retry policy
- Modem detection, pool management

Real-world test results (2026-07-23):
- V90 → 25153038: 4 pages, V.17 12000 bps, 2:33 ✅
- USR5637 → 25153038: 4 pages, V.17 12000 bps, 2:26 ✅
- Both faxes confirmed received on remote machine

Tested: loopback (100% pixel match), multi-page, all input formats,
cover pages, OCR verify, API endpoints, worker processing.
13 unit tests pass, 0 new clippy warnings.
2026-07-24 18:47:15 +08:00

14 KiB

OCR 优化建议

🎯 优化目标

当前问题:

  • English OCR: ✅ 99%准确度(完美)
  • Chinese OCR: ⚠️ 60%准确度(需改进)
  • Processing time: 可优化空间

目标:

  • Chinese OCR准确度: 60% → 85%+
  • Processing speed: 保持或提升
  • 商业部署: 生产就绪

📊 优化方案对比

方案 1: 安装更好的字体数据包(推荐)⭐⭐⭐⭐⭐

优势:

✅ 最简单
✅ 免费
✅ 快速提升准确度
✅ 无需代码修改

实施:

# 1. 下载最佳语言数据包
curl -L -o /opt/homebrew/share/tessdata/chi_tra_best.traineddata \
  https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_tra.traineddata

curl -L -o /opt/homebrew/share/tessdata/chi_sim_best.traineddata \
  https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_sim.traineddata

curl -L -o /opt/homebrew/share/tessdata/jpn_best.traineddata \
  https://github.com/tesseract-ocr/tessdata_best/raw/main/jpn.traineddata

# 2. 使用最佳数据包
tesseract input.png stdout -l chi_tra_best

# 3. 验证改进
tesseract --list-langs

效果预估:

Chinese Traditional: 60% → 85%
Chinese Simplified:  60% → 85%
Japanese:           60% → 85%

成本:

下载大小: ~100MB per language
加载时间: +200ms
内存占用: +150MB

方案 2: 图像预处理(高性价比)⭐⭐⭐⭐

优势:

✅ 提升所有语言准确度
✅ Rust实现
✅ 无额外依赖

实施:

// src/ocr/preprocess.rs
use image::{ImageBuffer, Luma};

pub struct ImagePreprocessor {
    target_dpi: u32,
}

impl ImagePreprocessor {
    pub fn preprocess_for_ocr(image: &[u8]) -> Result<Vec<u8>> {
        let img = image::load_from_memory(image)?;
        
        // 1. 提高分辨率
        let img = img.resize_exact(204 * 3, 196 * 3, image::imageops::FilterType::Lanczos3);
        
        // 2. 灰度化
        let img = img.grayscale();
        
        // 3. 二值化
        let img = self.binarize(&img);
        
        // 4. 噪点去除
        let img = self.remove_noise(&img);
        
        // 5. 边缘增强
        let img = self.enhance_edges(&img);
        
        Ok(img.to_bytes())
    }
    
    fn binarize(img: &DynamicImage) -> DynamicImage {
        // Adaptive thresholding
        let threshold = 128;
        img.brighten(20).contrast(1.2)
    }
    
    fn remove_noise(img: &DynamicImage) -> DynamicImage {
        // Apply Gaussian blur for noise removal
        img.blur(1.0)
    }
    
    fn enhance_edges(img: &DynamicImage) -> DynamicImage {
        // Sharpen edges for better text recognition
        img.sharpen(3.0)
    }
}

集成:

// 在OCR处理前预处理
let preprocessed = ImagePreprocessor::preprocess_for_ocr(&raw_data)?;
let ocr_result = ocr_processor.process_data(&preprocessed, "tiff")?;

效果预估:

All languages: +10-15% accuracy
Processing: +50ms preprocessing
Quality: High

方案 3: 参数优化(低成本)⭐⭐⭐⭐⭐

优势:

✅ 免费
✅ 快速
✅ 无需额外安装

实施:

// src/ocr/mod.rs
impl OcrProcessor {
    pub fn new_optimized() -> Self {
        Self {
            config: OcrConfig {
                // 使用 LSTM 神经网络引擎(最准确)
                oem: OcrEngineMode::NeuralNetLstmOnly,
                
                // 自动页面分割(最灵活)
                psm: PageSegMode::Auto,
                
                // 高DPI传真图像
                dpi: 204,  // 或更高: 300
                
                // 多语言组合
                language: OcrLanguage::Multi(vec!["chi_tra", "eng"]),
            },
        }
    }
    
    pub fn process_with_optimization(&self, image_path: &Path) -> Result<OcrResult> {
        let mut args = vec![
            image_path.to_string_lossy().to_string(),
            "stdout".to_string(),
            
            // 最佳参数
            "-l".to_string(), self.get_best_language_combo(),
            "--dpi".to_string(), "204".to_string(),
            "--oem".to_string(), "1".to_string(),  // LSTM only
            "--psm".to_string(), "3".to_string(),  // Auto
            
            // 高级参数
            "--dpi".to_string(), "300".to_string(),  // 提高DPI
            "quiet".to_string(),
        ];
        
        // 添加语言特定配置
        args.push("-c".to_string());
        args.push("tessedit_char_whitelist=0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz".to_string());
        
        // 执行OCR
        let output = Command::new(&self.tesseract_path)
            .args(&args)
            .output()?;
        
        Ok(self.parse_result(output))
    }
    
    fn get_best_language_combo(&self) -> String {
        match &self.config.language {
            OcrLanguage::ChineseTraditional => "chi_tra_best+eng",
            OcrLanguage::ChineseSimplified => "chi_sim_best+eng",
            OcrLanguage::Japanese => "jpn_best+eng",
            _ => "eng",
        }
    }
}

最佳参数:

# Tesseract参数优化
--oem 1           # LSTM神经网络(最准确)
--psm 3           # 自动页面分割(最灵活)
--dpi 300         # 高分辨率
-l chi_tra_best   # 最佳语言包

效果预估:

Chinese: +5-10% accuracy
Free: Yes
Time: +50ms

方案 4: 字体优化(核心)⭐⭐⭐⭐⭐

问题根源:

中文OCR准确度低的原因:
1. 测试图片使用Helvetica字体(不支持中文)
2. PIL默认字体不支持中文字符
3. 需要使用专门的中文字体

解决方案:

# scripts/cover_multilang.py
def create_chinese_cover_with_font(text, output_path):
    from PIL import Image, ImageDraw, ImageFont
    
    W, H = 1728, 2291
    img = Image.new('L', (W, H), 255)
    draw = ImageDraw.Draw(img)
    
    # 使用中文字体(关键)
    CHINESE_FONT_PATHS = [
        '/System/Library/AssetsV2/com_apple_MobileAsset_Font8/86ba2c91f017a3749571a82f2c6d890ac7ffb2fb.asset/AssetData/PingFang.ttc',
        '/System/Library/Fonts/STHeiti Medium.ttc',
        '/usr/share/fonts/truetype/wqy/wqy-zenhei.ttc',  # Linux
        '/usr/share/fonts/opentype/noto/NotoSansCJK-Regular.ttc',  # Linux
    ]
    
    font_path = None
    for path in CHINESE_FONT_PATHS:
        if os.path.exists(path):
            font_path = path
            break
    
    if font_path:
        if 'PingFang' in font_path:
            # PingFang字体索引(关键)
            font_title = ImageFont.truetype(font_path, 72, index=10)  # TC-Semibold
            font_bold = ImageFont.truetype(font_path, 44, index=6)    # TC-Medium
            font_normal = ImageFont.truetype(font_path, 44, index=2)  # TC-Light
        else:
            font_title = ImageFont.truetype(font_path, 72)
            font_bold = ImageFont.truetype(font_path, 44)
            font_normal = ImageFont.truetype(font_path, 44)
    else:
        # 下载并安装字体(可选)
        print("Warning: No Chinese font found. Downloading...")
        # 可以在这里添加自动下载逻辑
    
    # 绘制中文文本
    lines = text.split('\n')
    y = 100
    for line in lines:
        draw.text((100, y), line, fill=0, font=font_normal)
        y += 60
    
    img.save(output_path, dpi=(204, 196))

效果预估:

Chinese OCR: 60% → 90%+ (字体匹配)
Key insight: OCR准确度取决于字体匹配度

方案 5: 多次OCR + 结果融合(高准确度)⭐⭐⭐

原理:

多次OCR处理不同参数,融合结果:
1. English only OCR
2. Chinese only OCR
3. Combined OCR
4. 选择最佳结果

实施:

pub fn multi_pass_ocr(&self, image_path: &Path) -> Result<OcrResult> {
    let mut results = Vec::new();
    
    // Pass 1: English only
    let eng_result = self.process_with_language(image_path, "eng")?;
    results.push(eng_result);
    
    // Pass 2: Chinese Traditional
    let chi_result = self.process_with_language(image_path, "chi_tra")?;
    results.push(chi_result);
    
    // Pass 3: Combined
    let combined_result = self.process_with_language(image_path, "chi_tra+eng")?;
    results.push(combined_result);
    
    // 选择最佳结果(最高置信度)
    let best_result = results.iter()
        .max_by_key(|r| r.word_count)
        .unwrap();
    
    Ok(best_result.clone())
}

效果预估:

Accuracy: +10%
Processing: +300ms (3次处理)
Quality: High

方案 6: 机器学习后处理(高级)⭐⭐⭐

原理:

使用机器学习纠正OCR错误:
1. 收集常见错误样本
2. 训练纠错模型
3. 应用到OCR结果

实施:

// src/ocr/postprocess.rs
pub struct OcrPostProcessor {
    // 常见错误纠正表
    error_correction_map: HashMap<String, String>,
}

impl OcrPostProcessor {
    pub fn correct_ocr_text(&self, text: &str) -> String {
        let mut corrected = text.clone();
        
        // 常见OCR错误纠正
        let corrections = vec![
            ("O", "0"),  // 数字0识别为字母O
            ("l", "1"),  // 数字1识别为字母l
            ("S", "5"),  // 数字5识别为字母S
            ("B", "8"),  // 数字8识别为字母B
        ];
        
        for (wrong, correct) in corrections {
            // 根据上下文判断是否应该纠正
            corrected = self.smart_replace(&corrected, wrong, correct);
        }
        
        corrected
    }
    
    fn smart_replace(&self, text: &str, wrong: &str, correct: &str) -> String {
        // 智能替换:只在数字上下文中替换
        // 例如:电话号码中的O→0,但单词中的O保留
        let mut result = text.clone();
        
        for (i, c) in text.char_indices() {
            if c.to_string() == wrong {
                // 检查周围字符是否是数字
                let prev = text.chars().nth(i-1);
                let next = text.chars().nth(i+1);
                
                if prev.map(|p| p.is_digit(10)).unwrap_or(false) ||
                   next.map(|n| n.is_digit(10)).unwrap_or(false) {
                    result.replace_range(i..i+1, correct);
                }
            }
        }
        
        result
    }
}

效果预估:

Accuracy: +5-8%
Processing: +20ms
Complexity: Medium

📋 实施优先级

⭐⭐⭐⭐⭐ 最高优先级(立即实施)

1. 安装最佳语言数据包

# 免费且效果最好
curl -L -o /opt/homebrew/share/tessdata/chi_tra_best.traineddata \
  https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_tra.traineddata

2. 使用正确的中文字体

# 确保封面生成使用PingFang字体
font_path = '/System/Library/.../PingFang.ttc'
font = ImageFont.truetype(font_path, 44, index=10)

预期效果:

Chinese OCR: 60% → 90%
Free: Yes
Time: 1-2 hours

⭐⭐⭐⭐ 高优先级(本周实施)

3. 参数优化

// 使用最佳Tesseract参数
--oem 1 --psm 3 --dpi 300

4. 图像预处理

// 实现图像预处理模块
ImagePreprocessor::preprocess_for_ocr(&raw_data)

预期效果:

All languages: +10-15%
Time: 2-3 days

⭐⭐⭐ 中优先级(下月实施)

5. 多次OCR融合

// 多次处理取最佳结果
multi_pass_ocr(&image_path)

6. 后处理纠错

// 智能纠错常见OCR错误
OcrPostProcessor::correct_ocr_text(&text)

🎯 预期最终效果

优化后准确度

Language Current Optimized Improvement
English 99% 99.5% +0.5%
Chinese Traditional 60% 90% +30%
Chinese Simplified 60% 90% +30%
Japanese 60% 85% +25%
Multi-language 80% 92% +12%

优化后性能

Metric Current Optimized Impact
Processing 200ms 250ms +50ms
Memory 50MB 100MB +50MB
Quality Good Excellent ⭐⭐⭐⭐⭐

📝 实施步骤

Step 1: 立即实施(1小时)

#!/bin/bash
# scripts/install_best_ocr.sh

echo "Installing best Tesseract language packs..."

# 下载最佳语言包
curl -L -o /opt/homebrew/share/tessdata/chi_tra_best.traineddata \
  https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_tra.traineddata

curl -L -o /opt/homebrew/share/tessdata/chi_sim_best.traineddata \
  https://github.com/tesseract-ocr/tessdata_best/raw/main/chi_sim.traineddata

curl -L -o /opt/homebrew/share/tessdata/jpn_best.traineddata \
  https://github.com/tesseract-ocr/tessdata_best/raw/main/jpn.traineddata

# 测试改进
tesseract --list-langs

echo "Testing improved OCR..."
cd /tmp
python3 /Users/accusys/telfax/scripts/create_test_image.py test_chinese_optimized.png chinese
tesseract test_chinese_optimized.png stdout -l chi_tra_best

echo "Optimization complete!"

Step 2: 本周实施(2-3天)

// src/ocr/mod.rs - 添加优化配置
impl OcrProcessor {
    pub fn with_best_config() -> Self {
        Self {
            config: OcrConfig {
                language: OcrLanguage::Multi(vec!["chi_tra_best", "eng"]),
                dpi: 300,
                psm: PageSegMode::Auto,
                oem: OcrEngineMode::NeuralNetLstmOnly,
            },
        }
    }
}

// src/ocr/preprocess.rs - 添加预处理
pub fn preprocess_for_ocr(image: &[u8]) -> Result<Vec<u8>> {
    // 图像预处理提升准确度
    // ...
}

Step 3: 下月实施(1-2周)

// src/ocr/postprocess.rs - 后处理纠错
pub fn correct_ocr_errors(text: &str) -> String {
    // 智能纠错
    // ...
}

💡 总结

最优方案组合:

方案1 + 方案2 + 方案4 = 最佳效果
- 最佳语言包(免费)
- 图像预处理(Rust)
- 正确字体(关键)
= Chinese OCR: 90%+

实施建议:

  1. ⭐⭐⭐⭐⭐ 立即安装最佳语言包(1小时)
  2. ⭐⭐⭐⭐⭐ 使用正确中文字体(关键)
  3. ⭐⭐⭐⭐ 本周实现图像预处理(2-3天)
  4. ⭐⭐⭐ 下月添加后处理纠错(1-2周)

预期结果:

✅ Chinese OCR: 90%+ 准确度
✅ All languages: 提升10-15%
✅ 生产就绪
✅ 商业部署合规

优化建议完成!预期将中文OCR准确度提升至90%+ ✨