# teleocr_java
**Repository Path**: wuyuan/teleocr_java
## Basic Information
- **Project Name**: teleocr_java
- **Description**: springboot4.1.1 集成 teleocr
- **Primary Language**: Unknown
- **License**: AGPL-3.0
- **Default Branch**: master
- **Homepage**: None
- **GVP Project**: No
## Statistics
- **Stars**: 0
- **Forks**: 0
- **Created**: 2026-10-04
- **Last Updated**: 2026-10-04
## Categories & Tags
**Categories**: Uncategorized
**Tags**: None
## README
# TeleOCR Java / Spring Boot 4
Java 版 [stefanj0/TeleOCR-ONNX](https://huggingface.co/stefanj0/TeleOCR-ONNX)(Qwen2.5-VL-3B 蒸馏 OCR,默认 `q4f16`)推理实现。
`ort/teleocr_ort.py` 的 Java 移植:Pillow 双三次 resize → 归一化 → patchify → vision_encoder → embed_tokens → KV-cache 贪心解码,**不依赖 Python / transformers**。
## 模块
| 模块 | 说明 |
| --- | --- |
| `teleocr-spring-boot-starter` | 自动配置 + `TeleOcrClient`(Spring Boot 4,也可脱离 Spring 直接用) |
| `teleocr-api` | 可直接运行的 Spring Boot 4 REST 服务(上传 / base64 / 流式 / 任务列表 / 健康检查) |
环境:JDK 25+、Maven 3.9+、ONNX Runtime 1.30.0、DJL tokenizers 0.38.0。
## 构建
```bash
# GPU(默认,com.microsoft.onnxruntime:onnxruntime_gpu)
mvn -DskipTests install
mvn -P gpu -DskipTests install # 等价,显式写法
# CPU(com.microsoft.onnxruntime:onnxruntime,jar 体积小得多)
mvn -P cpu -DskipTests install
```
> GPU 包约 670MB,默认构建就是它;没有 NVIDIA 环境的机器用 `-P cpu`。
> 两者 API 完全一致——CPU 包启动 `provider: cuda` 时会自动回退 CPU 并打 WARN,不会崩。
`teleocr-api/target/teleocr-api.jar` 为可执行 fat jar。
## 快速开始(starter)
```xml
io.github.wuyuan2009123
teleocr-spring-boot-starter
1.0.0
```
```java
try (TeleOcrClient client = TeleOcrClient.open(new TeleOcrSettings().dtype("q4f16"))) {
OcrResult result = client.recognize(imageBytes); // 默认文本任务
System.out.println(result.getText());
System.out.printf("%d tok, ttft=%dms, total=%dms%n",
result.getGeneratedTokens(), result.getTimeToFirstTokenMs(), result.getTotalDurationMs());
// 表格 / 公式 / 版面
String latex = client.recognizeText(image, OcrOptions.of(OcrTask.FORMULA));
String otsl = client.recognizeText(image, OcrOptions.of(OcrTask.TABLE));
// 自定义提示词 + 流式
client.recognize(image, OcrOptions.ofPrompt("提取图中所有金额,每行一个")
.maxNewTokens(2048)
.tokenListener(delta -> System.out.print(delta)));
}
```
脱离 Spring 也可以直接 `TeleOcrClient.open(...)`;在 Spring Boot 中只需引入 starter,会自动创建单例 `TeleOcrClient`(`destroyMethod = "close"`)。
## 配置(`teleocr.*`)
| 属性 | 默认值 | 说明 |
| --- | --- | --- |
| `teleocr.enabled` | `true` | 是否创建客户端 Bean |
| `teleocr.model-path` | 空 | 本地模型目录(含 `config.json` + `tokenizer.json` + `onnx/`),非空时不再下载 |
| `teleocr.repo-id` | `stefanj0/TeleOCR-ONNX` | 缺失文件时从 HF 拉取的仓库 |
| `teleocr.dtype` | `q4f16` | `q4f16`(~1.8GB) / `fp16`(~2.9GB) / `fp32`(~5.7GB) |
| `teleocr.provider` | `cpu` | `cpu` / `cuda` / `dml` / `coreml` / `rocm` / `dnnl` |
| `teleocr.device-id` | `0` | `cuda`、`dml` 使用的设备序号 |
| `teleocr.intra-op-threads` | `0` | 0 = ONNX Runtime 默认 |
| `teleocr.download` | `true` | 允许自动下载 |
| `teleocr.download-dir` | 空 | 默认 `~/.cache/teleocr-onnx/` |
| `teleocr.hf-endpoint` | `https://huggingface.co` | 镜像站可设 `https://hf-mirror.com` |
| `teleocr.max-new-tokens` | `4096` | 单请求生成上限 |
| `teleocr.repetition-penalty` | 空 | 空 = 用 `generation_config.json` 的 1.05 |
| `teleocr.default-task` | `TEXT` | 默认任务 |
模型文件**下载成真实目录**(不是 HF 的 symlink 缓存),因为新版 ONNX Runtime 会拒绝位于 blob 目录的外部数据文件。
## GPU(CUDA)
`cuda` / `dml` 执行提供程序要求换成 GPU 包:
```bash
mvn -P gpu -DskipTests install
```
```yaml
teleocr:
provider: cuda # 或 dml(Windows DirectML)
device-id: 0
dtype: q4f16
```
或在启动参数里:`java -jar teleocr-api.jar --teleocr.provider=cuda --teleocr.device-id=0`。
- `-P gpu` 会把 `com.microsoft.onnxruntime:onnxruntime` 换成 `com.microsoft.onnxruntime:onnxruntime_gpu`(同版本 1.30.0),包体积约 670MB。
- 运行时需要 **CUDA 12.x + cuDNN 9.x**(ORT 1.30 要求);jar 里只带 `onnxruntime_providers_cuda.dll`,CUDA/cuDNN 动态库必须由系统提供。
- 若指定 `cuda` 但当前包不含该 EP 或依赖缺失(例如用 CPU 包启动、CUDA 版本过旧),会自动回退到 CPU 并打印 WARN 日志,服务不会起不来。
### Windows 上准备 CUDA 12 + cuDNN 9(系统级)
ORT 的 jar 不自带 CUDA/cuDNN。系统已有 CUDA 12.8+ / cuDNN 9 可直接跳过;否则不必装几 GB 的完整 toolkit,直接取 NVIDIA 官方 wheel 里的 DLL(免登录)装到系统目录即可(需管理员):
```powershell
# 1) 下载 NVIDIA 官方 wheel 并取出 bin 下的 DLL
$tmp = "$env:TEMP\cuda-wheels"
foreach ($p in "nvidia-cudnn-cu12","nvidia-cublas-cu12","nvidia-cuda-runtime-cu12","nvidia-cuda-nvrtc-cu12") {
$j = Invoke-RestMethod "https://pypi.org/pypi/$p/json"
$u = ($j.releases.($j.info.version) | Where-Object filename -like "*win_amd64*").url
curl.exe -sL -o "$tmp\$p.whl" $u
Expand-Archive -Force "$tmp\$p.whl" "$tmp\$p"
}
$bin = "C:\Program Files\NVIDIA\CUDA-Runtime-12.9\bin"
New-Item -ItemType Directory -Force -Path $bin | Out-Null
Get-ChildItem $tmp -Recurse -Filter *.dll | Where-Object FullName -match "\\bin\\" |
ForEach-Object { Copy-Item $_.FullName $bin -Force }
# 2) 加到「系统」PATH 最前面(不动 CUDA_PATH,不写 System32)
$p = [Environment]::GetEnvironmentVariable('Path','Machine')
[Environment]::SetEnvironmentVariable('Path', "$bin;$p", 'Machine')
# 3) 广播 WM_SETTINGCHANGE,让已运行的程序(含资源管理器)刷新环境,无需重启
Add-Type -Namespace Win32 -Name N -MemberDefinition `
'[DllImport("user32.dll",CharSet=CharSet.Auto)]public static extern IntPtr SendMessageTimeout(IntPtr h,uint m,UIntPtr w,string l,uint f,uint t,out UIntPtr r);'
[UIntPtr]$r = 0
[Win32.N]::SendMessageTimeout([IntPtr]0xffff,0x001A,[UIntPtr]::Zero,"Environment",2,5000,[ref]$r) | Out-Null
```
安全要点:
- 不改动 `CUDA_PATH` / 不覆盖 `C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v11.8`,原有 CUDA 11.8(`cudart64_110`)照旧可用;
- 不复制到 `C:\Windows\System32`,避免与系统里已有的 cuDNN 8 抢解析顺序;
- `C:\llama_server` 之类自带 CUDA DLL 的程序仍优先加载自己目录下的版本(Windows 先查程序所在目录);
- 卸载:删掉 PATH 里那一条 + 删除该目录即可。
> 如果本机还装着旧的 CUDA 11.8:可以一并卸载(Nsight 用 `MsiExec /X{GUID}`,CUDA 组件用
> `"C:\Program Files\NVIDIA Corporation\Installer2\InstallerCore\NVI2.DLL",UninstallPackage `),
> 再删掉 `CUDA_PATH` / `CUDA_PATH_V11_8` / `NVTOOLSEXT_PATH`。卸载器不会删手动拷进去的 cuDNN,
> `...\CUDA\v11.8` 目录需手动清理。注意卸载后**没有 nvcc**,需要编译 CUDA 代码时请装完整 CUDA Toolkit 12.x。
装好后直接运行,不需要任何 PATH 前缀:
```powershell
java -Xmx3g -jar teleocr-api\target\teleocr-api.jar --teleocr.provider=cuda
```
### 本机实测(GTX 1650 4GB)
实测同一张图(960×320,26 个生成 token):
| provider | TTFT | 总耗时 | tokens/s | 显存 |
| --- | --- | --- | --- | --- |
| cpu | 57.9 s | 61.0 s | 8.1 | — |
| cuda (GTX 1650) | 7.4 s | 8.8 s | 17.3 | 2.9 / 4.0 GB |
`q4f16` 三个图共约 1.8GB,4GB 显存能放下;视觉 token 更多的大图(`LAYOUT` 任务会缩放到 1036×1036)显存会继续涨,必要时减小图片尺寸或改用 `fp16`/CPU。
## REST API
```bash
java -jar teleocr-api/target/teleocr-api.jar
# http://localhost:8080/
```
| 方法 | 路径 | 说明 |
| --- | --- | --- |
| `POST` | `/api/v1/ocr` | multipart `file`,参数 `task`/`prompt`/`maxNewTokens`/`repetitionPenalty`,返回 JSON |
| `POST` | `/api/v1/ocr/stream` | 同上,`text/plain` 流式输出 |
| `POST` | `/api/v1/ocr/base64` | JSON `{image, task, prompt, maxNewTokens, repetitionPenalty}` |
| `GET` | `/api/v1/ocr/tasks` | 任务名 → 提示词 |
| `GET` | `/api/v1/ocr/health` | 状态 + 解析后的 dtype/provider/deviceId |
```bash
curl -F "file=@page.png" -F "task=TEXT" http://localhost:8080/api/v1/ocr
curl -F "file=@formula.png" -F "task=FORMULA" http://localhost:8080/api/v1/ocr/stream
```
任务:`TEXT`、`TABLE`(OTSL)、`FORMULA`(LaTeX)、`CODE`、`LAYOUT`、`LAYOUT_DISTORTED`(这两个会先把页面按 Pillow 双三次缩放到 1036×1036)、`FIGURE`。
## 发布到 Maven Central
坐标:`io.github.wuyuan2009123:teleocr-spring-boot-starter:1.0.0`。
`teleocr-parent`(pom)会一起发布,`teleocr-api` 是示例服务,已排除在发布之外。
### 一次性准备
1. 用 GitHub 账号(`wuyuan2009123`)登录 ;`io.github.wuyuan2009123` 这个 namespace 会自动验证通过。
2. Account → **Generate User Token**,把用户名/密码写进 `~/.m2/settings.xml`(见 `settings.xml.example`)。
3. 生成 GPG 密钥并把公钥推到公共 keyserver:
```bash
gpg --full-generate-key # RSA 4096,填与 developer email 一致
gpg --list-secret-keys --keyid-format=long # 记下 sec 行 rsa4096/
gpg --keyserver keyserver.ubuntu.com --send-keys
```
用 `-Dgpg.passphrase=...` 免交互签名时,需在 `~/.gnupg/gpg-agent.conf` 加一行 `allow-loopback-pinentry`,然后 `gpgconf --kill gpg-agent`。
### 发布
```bash
mvn -Prelease clean deploy
```
`release` profile 会产出 `*-sources.jar`、`*-javadoc.jar` 和每个文件的 `.asc` 签名,再由 `central-publishing-maven-plugin` 打包上传:
| 插件 | 作用 |
| --- | --- |
| `maven-source-plugin` 3.4.0 | sources jar |
| `maven-javadoc-plugin` 3.12.0 | javadoc jar(`doclint=none`,离线可构建) |
| `maven-gpg-plugin` 3.2.8 | GPG 签名(绑定 `verify` 阶段) |
| `central-publishing-maven-plugin` 0.11.0 | 上传 + 发布 |
默认 `autoPublish=true` + `waitUntil=published`:校验通过后直接发布。想先在网页上人工确认,改成 `mvn -Prelease deploy -DautoPublish=false`,再到 点 Publish。
**发布后不可删除、不可覆盖**,出错只能升版本号重发。
### 只验证打包、不上传
```bash
mvn -Prelease clean package -DskipTests # 生成 sources/javadoc jar,不签名不上传
```
### 升版本
```bash
mvn versions:set -DnewVersion=1.0.1 -DprocessAllModules
mvn -Prelease clean deploy
```
### 注意
- 许可证是 **AGPL-3.0**(仓库根目录 `LICENSE`),下游使用会触发 copyleft 义务;若要放宽请一并改 `LICENSE` 和两个 POM 的 ``。
- starter 默认依赖 `onnxruntime_gpu`(约 670MB)。消费方想用 CPU 包:
```xml
io.github.wuyuan2009123
teleocr-spring-boot-starter
1.0.0
com.microsoft.onnxruntime
onnxruntime_gpu
com.microsoft.onnxruntime
onnxruntime
1.30.0
```
## 实现要点
- **Pillow 双三次重采样**:按 `js/pil_resize.js` 逐行移植 `libImaging/Resample.c` 的定点系数计算,避免 Java `Graphics2D` 缩放带来的像素差异。
- **预处理**:`smart_resize`(28 的倍数 + min/max_pixels)→ 归一化(rescale 1/255、mean/std)→ patchify 成 `[patches, 1176]`。
- **MRoPE**:按 `<|image_pad|>` 段落展开 `[3, 1, S]` 位置 id,与 Python 参考一致。
- **解码**:视觉特征替换 `<|image_pad|>` 位置的 embedding;`past_key_values.*` / `present.*` 复用 KV cache,每步只喂 1 个 token;repetition penalty 作用在已出现 token 上,argmax 采样。
- **分词**:直接用 HF 的 `tokenizer.json`(DJL 绑定 Rust `tokenizers`),与 Python 端完全一致。
- **线程**:`TeleOcrEngine.generate` 串行执行(单步会产出 `[1, seq, 151936]` 的 logits 张量,并发会打爆内存)。