# teleocr_java **Repository Path**: wuyuan/teleocr_java ## Basic Information - **Project Name**: teleocr_java - **Description**: springboot4.1.1 集成 teleocr - **Primary Language**: Unknown - **License**: AGPL-3.0 - **Default Branch**: master - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-10-04 - **Last Updated**: 2026-10-04 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # TeleOCR Java / Spring Boot 4 Java 版 [stefanj0/TeleOCR-ONNX](https://huggingface.co/stefanj0/TeleOCR-ONNX)(Qwen2.5-VL-3B 蒸馏 OCR,默认 `q4f16`)推理实现。 `ort/teleocr_ort.py` 的 Java 移植:Pillow 双三次 resize → 归一化 → patchify → vision_encoder → embed_tokens → KV-cache 贪心解码,**不依赖 Python / transformers**。 ## 模块 | 模块 | 说明 | | --- | --- | | `teleocr-spring-boot-starter` | 自动配置 + `TeleOcrClient`(Spring Boot 4,也可脱离 Spring 直接用) | | `teleocr-api` | 可直接运行的 Spring Boot 4 REST 服务(上传 / base64 / 流式 / 任务列表 / 健康检查) | 环境:JDK 25+、Maven 3.9+、ONNX Runtime 1.30.0、DJL tokenizers 0.38.0。 ## 构建 ```bash # GPU(默认,com.microsoft.onnxruntime:onnxruntime_gpu) mvn -DskipTests install mvn -P gpu -DskipTests install # 等价,显式写法 # CPU(com.microsoft.onnxruntime:onnxruntime,jar 体积小得多) mvn -P cpu -DskipTests install ``` > GPU 包约 670MB,默认构建就是它;没有 NVIDIA 环境的机器用 `-P cpu`。 > 两者 API 完全一致——CPU 包启动 `provider: cuda` 时会自动回退 CPU 并打 WARN,不会崩。 `teleocr-api/target/teleocr-api.jar` 为可执行 fat jar。 ## 快速开始(starter) ```xml io.github.wuyuan2009123 teleocr-spring-boot-starter 1.0.0 ``` ```java try (TeleOcrClient client = TeleOcrClient.open(new TeleOcrSettings().dtype("q4f16"))) { OcrResult result = client.recognize(imageBytes); // 默认文本任务 System.out.println(result.getText()); System.out.printf("%d tok, ttft=%dms, total=%dms%n", result.getGeneratedTokens(), result.getTimeToFirstTokenMs(), result.getTotalDurationMs()); // 表格 / 公式 / 版面 String latex = client.recognizeText(image, OcrOptions.of(OcrTask.FORMULA)); String otsl = client.recognizeText(image, OcrOptions.of(OcrTask.TABLE)); // 自定义提示词 + 流式 client.recognize(image, OcrOptions.ofPrompt("提取图中所有金额,每行一个") .maxNewTokens(2048) .tokenListener(delta -> System.out.print(delta))); } ``` 脱离 Spring 也可以直接 `TeleOcrClient.open(...)`;在 Spring Boot 中只需引入 starter,会自动创建单例 `TeleOcrClient`(`destroyMethod = "close"`)。 ## 配置(`teleocr.*`) | 属性 | 默认值 | 说明 | | --- | --- | --- | | `teleocr.enabled` | `true` | 是否创建客户端 Bean | | `teleocr.model-path` | 空 | 本地模型目录(含 `config.json` + `tokenizer.json` + `onnx/`),非空时不再下载 | | `teleocr.repo-id` | `stefanj0/TeleOCR-ONNX` | 缺失文件时从 HF 拉取的仓库 | | `teleocr.dtype` | `q4f16` | `q4f16`(~1.8GB) / `fp16`(~2.9GB) / `fp32`(~5.7GB) | | `teleocr.provider` | `cpu` | `cpu` / `cuda` / `dml` / `coreml` / `rocm` / `dnnl` | | `teleocr.device-id` | `0` | `cuda`、`dml` 使用的设备序号 | | `teleocr.intra-op-threads` | `0` | 0 = ONNX Runtime 默认 | | `teleocr.download` | `true` | 允许自动下载 | | `teleocr.download-dir` | 空 | 默认 `~/.cache/teleocr-onnx/` | | `teleocr.hf-endpoint` | `https://huggingface.co` | 镜像站可设 `https://hf-mirror.com` | | `teleocr.max-new-tokens` | `4096` | 单请求生成上限 | | `teleocr.repetition-penalty` | 空 | 空 = 用 `generation_config.json` 的 1.05 | | `teleocr.default-task` | `TEXT` | 默认任务 | 模型文件**下载成真实目录**(不是 HF 的 symlink 缓存),因为新版 ONNX Runtime 会拒绝位于 blob 目录的外部数据文件。 ## GPU(CUDA) `cuda` / `dml` 执行提供程序要求换成 GPU 包: ```bash mvn -P gpu -DskipTests install ``` ```yaml teleocr: provider: cuda # 或 dml(Windows DirectML) device-id: 0 dtype: q4f16 ``` 或在启动参数里:`java -jar teleocr-api.jar --teleocr.provider=cuda --teleocr.device-id=0`。 - `-P gpu` 会把 `com.microsoft.onnxruntime:onnxruntime` 换成 `com.microsoft.onnxruntime:onnxruntime_gpu`(同版本 1.30.0),包体积约 670MB。 - 运行时需要 **CUDA 12.x + cuDNN 9.x**(ORT 1.30 要求);jar 里只带 `onnxruntime_providers_cuda.dll`,CUDA/cuDNN 动态库必须由系统提供。 - 若指定 `cuda` 但当前包不含该 EP 或依赖缺失(例如用 CPU 包启动、CUDA 版本过旧),会自动回退到 CPU 并打印 WARN 日志,服务不会起不来。 ### Windows 上准备 CUDA 12 + cuDNN 9(系统级) ORT 的 jar 不自带 CUDA/cuDNN。系统已有 CUDA 12.8+ / cuDNN 9 可直接跳过;否则不必装几 GB 的完整 toolkit,直接取 NVIDIA 官方 wheel 里的 DLL(免登录)装到系统目录即可(需管理员): ```powershell # 1) 下载 NVIDIA 官方 wheel 并取出 bin 下的 DLL $tmp = "$env:TEMP\cuda-wheels" foreach ($p in "nvidia-cudnn-cu12","nvidia-cublas-cu12","nvidia-cuda-runtime-cu12","nvidia-cuda-nvrtc-cu12") { $j = Invoke-RestMethod "https://pypi.org/pypi/$p/json" $u = ($j.releases.($j.info.version) | Where-Object filename -like "*win_amd64*").url curl.exe -sL -o "$tmp\$p.whl" $u Expand-Archive -Force "$tmp\$p.whl" "$tmp\$p" } $bin = "C:\Program Files\NVIDIA\CUDA-Runtime-12.9\bin" New-Item -ItemType Directory -Force -Path $bin | Out-Null Get-ChildItem $tmp -Recurse -Filter *.dll | Where-Object FullName -match "\\bin\\" | ForEach-Object { Copy-Item $_.FullName $bin -Force } # 2) 加到「系统」PATH 最前面(不动 CUDA_PATH,不写 System32) $p = [Environment]::GetEnvironmentVariable('Path','Machine') [Environment]::SetEnvironmentVariable('Path', "$bin;$p", 'Machine') # 3) 广播 WM_SETTINGCHANGE,让已运行的程序(含资源管理器)刷新环境,无需重启 Add-Type -Namespace Win32 -Name N -MemberDefinition ` '[DllImport("user32.dll",CharSet=CharSet.Auto)]public static extern IntPtr SendMessageTimeout(IntPtr h,uint m,UIntPtr w,string l,uint f,uint t,out UIntPtr r);' [UIntPtr]$r = 0 [Win32.N]::SendMessageTimeout([IntPtr]0xffff,0x001A,[UIntPtr]::Zero,"Environment",2,5000,[ref]$r) | Out-Null ``` 安全要点: - 不改动 `CUDA_PATH` / 不覆盖 `C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v11.8`,原有 CUDA 11.8(`cudart64_110`)照旧可用; - 不复制到 `C:\Windows\System32`,避免与系统里已有的 cuDNN 8 抢解析顺序; - `C:\llama_server` 之类自带 CUDA DLL 的程序仍优先加载自己目录下的版本(Windows 先查程序所在目录); - 卸载:删掉 PATH 里那一条 + 删除该目录即可。 > 如果本机还装着旧的 CUDA 11.8:可以一并卸载(Nsight 用 `MsiExec /X{GUID}`,CUDA 组件用 > `"C:\Program Files\NVIDIA Corporation\Installer2\InstallerCore\NVI2.DLL",UninstallPackage `), > 再删掉 `CUDA_PATH` / `CUDA_PATH_V11_8` / `NVTOOLSEXT_PATH`。卸载器不会删手动拷进去的 cuDNN, > `...\CUDA\v11.8` 目录需手动清理。注意卸载后**没有 nvcc**,需要编译 CUDA 代码时请装完整 CUDA Toolkit 12.x。 装好后直接运行,不需要任何 PATH 前缀: ```powershell java -Xmx3g -jar teleocr-api\target\teleocr-api.jar --teleocr.provider=cuda ``` ### 本机实测(GTX 1650 4GB) 实测同一张图(960×320,26 个生成 token): | provider | TTFT | 总耗时 | tokens/s | 显存 | | --- | --- | --- | --- | --- | | cpu | 57.9 s | 61.0 s | 8.1 | — | | cuda (GTX 1650) | 7.4 s | 8.8 s | 17.3 | 2.9 / 4.0 GB | `q4f16` 三个图共约 1.8GB,4GB 显存能放下;视觉 token 更多的大图(`LAYOUT` 任务会缩放到 1036×1036)显存会继续涨,必要时减小图片尺寸或改用 `fp16`/CPU。 ## REST API ```bash java -jar teleocr-api/target/teleocr-api.jar # http://localhost:8080/ ``` | 方法 | 路径 | 说明 | | --- | --- | --- | | `POST` | `/api/v1/ocr` | multipart `file`,参数 `task`/`prompt`/`maxNewTokens`/`repetitionPenalty`,返回 JSON | | `POST` | `/api/v1/ocr/stream` | 同上,`text/plain` 流式输出 | | `POST` | `/api/v1/ocr/base64` | JSON `{image, task, prompt, maxNewTokens, repetitionPenalty}` | | `GET` | `/api/v1/ocr/tasks` | 任务名 → 提示词 | | `GET` | `/api/v1/ocr/health` | 状态 + 解析后的 dtype/provider/deviceId | ```bash curl -F "file=@page.png" -F "task=TEXT" http://localhost:8080/api/v1/ocr curl -F "file=@formula.png" -F "task=FORMULA" http://localhost:8080/api/v1/ocr/stream ``` 任务:`TEXT`、`TABLE`(OTSL)、`FORMULA`(LaTeX)、`CODE`、`LAYOUT`、`LAYOUT_DISTORTED`(这两个会先把页面按 Pillow 双三次缩放到 1036×1036)、`FIGURE`。 ## 发布到 Maven Central 坐标:`io.github.wuyuan2009123:teleocr-spring-boot-starter:1.0.0`。 `teleocr-parent`(pom)会一起发布,`teleocr-api` 是示例服务,已排除在发布之外。 ### 一次性准备 1. 用 GitHub 账号(`wuyuan2009123`)登录 ;`io.github.wuyuan2009123` 这个 namespace 会自动验证通过。 2. Account → **Generate User Token**,把用户名/密码写进 `~/.m2/settings.xml`(见 `settings.xml.example`)。 3. 生成 GPG 密钥并把公钥推到公共 keyserver: ```bash gpg --full-generate-key # RSA 4096,填与 developer email 一致 gpg --list-secret-keys --keyid-format=long # 记下 sec 行 rsa4096/ gpg --keyserver keyserver.ubuntu.com --send-keys ``` 用 `-Dgpg.passphrase=...` 免交互签名时,需在 `~/.gnupg/gpg-agent.conf` 加一行 `allow-loopback-pinentry`,然后 `gpgconf --kill gpg-agent`。 ### 发布 ```bash mvn -Prelease clean deploy ``` `release` profile 会产出 `*-sources.jar`、`*-javadoc.jar` 和每个文件的 `.asc` 签名,再由 `central-publishing-maven-plugin` 打包上传: | 插件 | 作用 | | --- | --- | | `maven-source-plugin` 3.4.0 | sources jar | | `maven-javadoc-plugin` 3.12.0 | javadoc jar(`doclint=none`,离线可构建) | | `maven-gpg-plugin` 3.2.8 | GPG 签名(绑定 `verify` 阶段) | | `central-publishing-maven-plugin` 0.11.0 | 上传 + 发布 | 默认 `autoPublish=true` + `waitUntil=published`:校验通过后直接发布。想先在网页上人工确认,改成 `mvn -Prelease deploy -DautoPublish=false`,再到 点 Publish。 **发布后不可删除、不可覆盖**,出错只能升版本号重发。 ### 只验证打包、不上传 ```bash mvn -Prelease clean package -DskipTests # 生成 sources/javadoc jar,不签名不上传 ``` ### 升版本 ```bash mvn versions:set -DnewVersion=1.0.1 -DprocessAllModules mvn -Prelease clean deploy ``` ### 注意 - 许可证是 **AGPL-3.0**(仓库根目录 `LICENSE`),下游使用会触发 copyleft 义务;若要放宽请一并改 `LICENSE` 和两个 POM 的 ``。 - starter 默认依赖 `onnxruntime_gpu`(约 670MB)。消费方想用 CPU 包: ```xml io.github.wuyuan2009123 teleocr-spring-boot-starter 1.0.0 com.microsoft.onnxruntime onnxruntime_gpu com.microsoft.onnxruntime onnxruntime 1.30.0 ``` ## 实现要点 - **Pillow 双三次重采样**:按 `js/pil_resize.js` 逐行移植 `libImaging/Resample.c` 的定点系数计算,避免 Java `Graphics2D` 缩放带来的像素差异。 - **预处理**:`smart_resize`(28 的倍数 + min/max_pixels)→ 归一化(rescale 1/255、mean/std)→ patchify 成 `[patches, 1176]`。 - **MRoPE**:按 `<|image_pad|>` 段落展开 `[3, 1, S]` 位置 id,与 Python 参考一致。 - **解码**:视觉特征替换 `<|image_pad|>` 位置的 embedding;`past_key_values.*` / `present.*` 复用 KV cache,每步只喂 1 个 token;repetition penalty 作用在已出现 token 上,argmax 采样。 - **分词**:直接用 HF 的 `tokenizer.json`(DJL 绑定 Rust `tokenizers`),与 Python 端完全一致。 - **线程**:`TeleOcrEngine.generate` 串行执行(单步会产出 `[1, seq, 151936]` 的 logits 张量,并发会打爆内存)。