# es **Repository Path**: middleware-lab/es ## Basic Information - **Project Name**: es - **Description**: No description available - **Primary Language**: Unknown - **License**: Not specified - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-09-15 - **Last Updated**: 2026-09-16 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # ES 模拟器 · Elasticsearch 场景实验室 用 **Go** 从零模拟 Elasticsearch 的核心行为(倒排索引、分词器、BM25 打分、聚合、分页限制……),并以 **React** 前端交互式演示常见的 ES 使用场景、特性,以及实际开发中最容易踩的坑。 > 本项目不依赖真实 ES 集群,后端就是一个用 Go 手写的"迷你 ES"。所有查询错误(如 `Result window is too large`、`Fielddata is disabled on text fields`、映射冲突)都与真实 ES 的报错形态一致。 ## 功能一览 后端内置 14 个场景,分为三类: ### 使用场景 | 场景 | 说明 | | --- | --- | | 全文检索入门:match | 分词 + 倒排索引 + BM25 相关性排序 | | term 与 match 的区别 | keyword 精确匹配 vs text 全文匹配 | | 布尔组合查询:bool | must / should / must_not / filter 组合 | | 范围查询:range | 数值与日期区间过滤 | | 深度分页:search_after | sort + 游标的正确翻页方式 | ### 特性 | 场景 | 说明 | | --- | --- | | 短语匹配:match_phrase | 词项位置 + slop 容忍间隔 | | 模糊搜索:fuzzy | 编辑距离容错(AUTO 档位) | | 聚合分析 | terms 分桶、avg 等指标、date_histogram 时间直方图 | | 相关性打分 | BM25 的 TF / IDF / 长度归一化与 `_explain` | | 同义词搜索 | synonym 分析器双向展开同义词 | ### 常见问题(真实报错复现) | 场景 | 报错 | | --- | --- | | 深分页报错 | `Result window is too large, from + size must be less than or equal to: [10000]` | | 分词器不匹配 | 搜 `runs` 找不到 `running`(standard 与 english 分析器对比) | | text 字段排序/聚合 | `Fielddata is disabled on text fields by default...` | | 映射类型冲突 | `mapper [price] cannot be changed from type [long] to [text]` | 另有独立的**分词演练场**:同一段文本分别交给 standard / english / whitespace / keyword / synonym 五种分析器,直观展示 token 差异。 ## 项目结构 ``` es/ ├── backend/ # Go 后端(仅标准库,零第三方依赖) │ ├── main.go # 入口:HTTP 服务 + 可选托管前端静态资源 │ └── internal/ │ ├── engine/ # 迷你 ES 引擎 │ │ ├── analyzer.go # 分词器:standard / english / whitespace / keyword / synonym │ │ ├── engine.go # 索引、映射、倒排索引、动态映射类型推断 │ │ ├── query.go # 查询 DSL 解析与执行 + BM25 打分 │ │ ├── aggregate.go # terms / metric / date_histogram 聚合 │ │ ├── search.go # 检索流水线:排序、分页、search_after、高亮 │ │ ├── errors.go # ES 风格错误(search_phase / mapper_parsing ...) │ │ └── util.go │ ├── scenarios/ # 14 个内置场景(数据 + 默认查询 + 讲解文案) │ └── api/ # HTTP API + CORS + 静态资源 └── frontend/ # React 18 + Vite + TypeScript 前端 └── src/ ├── components/ # 场景列表 / 讲解面板 / DSL 编辑器 / 结果与图表 / 分词演练场 ├── api.ts # 后端 API 封装 └── types.ts # 与后端 JSON 对齐的类型 ``` ## 快速开始 ### 方式一:一键运行(后端托管前端) ```bash # 1. 构建前端(首次需要 npm install) cd frontend npm install npm run build # 2. 启动后端(自动托管 ../frontend/dist) cd ../backend go run . # => http://localhost:9200 ``` 浏览器打开 即可使用。端口可通过 `PORT` 环境变量修改(8080 常被本机其他服务占用,建议 9200——正好也是 ES 的默认端口)。 ### 方式二:前端开发模式(热更新) ```bash # 终端 1 cd backend && PORT=9200 go run . # 终端 2 cd frontend && npm run dev # => http://localhost:5173 ``` > 若后端不在 `localhost:9200`,可通过环境变量指定:`VITE_API_BASE=http://localhost:9200 npm run dev`。 ### 依赖与镜像 - Go 1.23+:后端只使用标准库,无需下载任何第三方模块。 - Node 18+:前端依赖默认走本机已配置的 npm 镜像(`registry.npmmirror.com`),无需额外配置。 ## HTTP API | 方法 | 路径 | 说明 | | --- | --- | --- | | GET | `/api/health` | 健康检查 | | GET | `/api/scenarios` | 场景列表(含默认查询) | | GET | `/api/scenarios/{id}` | 场景详情(映射 + 文档) | | POST | `/api/scenarios/{id}/search` | 执行查询(请求体即查询 DSL) | | POST | `/api/scenarios/{id}/reset` | 重置场景索引缓存 | | GET | `/api/analyzers` | 分析器列表 | | GET | `/api/analyze?text=...&analyzer=...` | 分词结果 | 查询 DSL 示例(与 ES 语法一致的子集): ```json { "query": { "bool": { "must": [{ "match": { "description": "mouse" } }], "filter": [{ "range": { "price": { "gte": 100, "lte": 500 } } }], "must_not": [{ "term": { "brand": "Razer" } }] } }, "from": 0, "size": 10, "sort": [{ "price": "asc" }], "aggs": { "by_category": { "terms": { "field": "category" } } }, "highlight": { "enabled": true }, "explain": true } ``` ## 引擎实现说明(模拟边界) - **分词**:standard 按非字母数字切分并小写;english 额外做简化版词干提取(`runs→run`、`running→run`);whitespace 保留标点;keyword 整体成一个词项;synonym 在两侧展开同义词表。 - **打分**:简化版 BM25(k1=1.2, b=0.75),IDF 取 `ln(1 + (N-df+0.5)/(df+0.5))`。 - **位置信息**:倒排索引记录词项位置,支撑 match_phrase 与 slop。 - **动态映射**:按首条文档推断字段类型,类型冲突即抛 `mapper_parsing_exception`(复现真实 ES 行为)。 - **max_result_window**:默认 10000;深分页场景将其调小为 10,用少量数据即可复现真实报错。 - 中文按单字切分(模拟 standard 对未分词 CJK 文本的处理),生产环境请使用 IK 等中文分词器。