Files
fscan/webscan/fingerprint/enhanced_test.go
T
ZacharyZcR 760c8ea502 v2.1.2 核心优化与多架构发布 (#561)
* feat: v2.1.0 核心重构与功能增强

## 架构重构
- 全局变量消除,迁移至 Config/State 对象
- SMB 插件融合(smb/smb2/smbghost/smbinfo)
- 服务探测重构,实现 Nmap 风格 fallback 机制
- 输出系统重构,TXT 实时刷盘 + 双写机制
- i18n 框架升级至 go-i18n

## 性能优化
- 正则表达式预编译
- 内存优化 map[string]struct{}
- 并发指纹匹配
- SOCKS5 连接复用
- 滑动窗口调度 + 自适应线程池

## 新功能
- Web 管理界面
- 多格式 POC 适配(xray/afrog)
- 增强指纹库(3139条)
- Favicon hash 指纹识别
- 插件选择性编译(Build Tags)
- fscan-lab 靶场环境
- 默认端口扩展(62→133)

## 构建系统
- 添加 no_local tag 支持排除本地插件
- 多版本构建:fscan/fscan-nolocal/fscan-web
- CI 添加 snapshot 模式支持仅测试构建

## Bug 修复
- 修复 120+ 个问题,包括 RDP panic、批量扫描漏报、
  JSON 输出格式、Redis 检测、Context 超时等

## 测试增强
- 单元测试覆盖率 74-100%
- 并发安全测试
- 集成测试(Web/端口/服务/SSH/ICMP)

* fix(ci): 移除 PR 对 Project 自动化的触发

* fix: Elasticsearch未授权检测优先于爆破 (#554)

* fix: 修复RDP爆破高误报率问题 (#555)

- 移除 screen.go 中错误的认证结果覆盖逻辑
- 启用 NLA 协议的 ErrorCode 字段检测
- 添加 PubKeyAuth 验证确保认证真正成功
- 修复 io.go 中错误被静默忽略的问题
- 修复 socket.go/io.go 中可能导致 panic 的代码
- 修复 screen.go 中文件句柄泄漏和 log.Panic

* fix: 修复-user/-pwd凭据参数不生效的问题

问题原因:
- Parse()解析凭据后更新globalConfig
- 但BuildConfigFromFlags()创建新Config时使用默认字典
- 导致解析的UserPassPairs等凭据信息被丢弃

修复内容:
1. initialize.go: 将Parse解析的凭据结果应用到新Config
2. credential.go: 单用户密码对时创建UserPassPairs
3. rdp.go: 单凭据测试时跳过指纹识别,减少连接次数

* feat: RDP使用NLA仅验证模式,避免挤掉已登录用户

- 添加ErrNLAAuthSuccess标志用于NLA验证成功信号
- tpkt层支持nlaAuthOnly模式,验证成功后不建立完整会话
- x224层正确传播NLA验证结果
- rdpCrack改用NlaAuth进行凭据验证

* fix: 修复进度条在Windows终端满屏重复输出的问题

- 添加终端宽度检测,动态调整进度条长度
- 使用空格覆盖清除旧内容,避免残留
- 简化进度条格式,确保不超过终端宽度

* feat: 优化日志颜色方案,区分漏洞和普通信息

- 新增 LogVuln 级别(红色),用于漏洞和重要发现
- 密码爆破成功、未授权访问、POC漏洞等改用红色显示
- 普通信息(扫描统计等)改为白色
- Web指纹保持绿色

* refactor: 精简化输出,移除冗余启动信息

- 移除showParseSummary开局配置输出
- 移除LogPluginInfo/LogPluginInfoWithPort插件信息输出
- 移除alive_scanner冗余统计输出
- 移除port_scan_start扫描开始提示
- 移除handleUDPPorts SNMP死代码
- 移除相关i18n条目

* chore: 版本号更新为2.1.1

* fix: 降级依赖版本以保持Go 1.20兼容性

* feat(ldap): 添加NTLM Hash认证支持 (#433)

* chore: 清理无用的 replace 指令

* fix(ping): 修复 TTL expired 导致主机误判为存活的问题

在 ExecCommandPing 中增加错误关键词检测,当 ping 输出包含
TTL expired、Destination unreachable 等错误信息时,不再将
目标主机标记为存活。

Fixes #454

* fix(proxy): 修复透明代理导致输出全端口的问题

在代理初始化时主动探测代理行为,通过连接 RFC 5737 保留的
测试地址来检测是否存在"全回显"问题。如果探测到代理不可靠,
则在端口扫描时跳过所有端口,避免误报。

- 新增 proxyReliable 标志位标记代理可靠性
- 新增 ProbeProxyBehavior 函数探测代理行为
- 端口扫描前检查代理可靠性并输出警告

Fixes #495

* refactor: 移动debug模块到common/debug子包

* fix(web): 修复-u模式下Web插件未执行的问题

* fix: 优化输出格式和颜色显示

- 网段统计格式改为 10.253.0.0/16 网段存活: 26
- WebTitle基础信息改为白色,指纹识别单独绿色输出
- 移除重复的端口数量输出

* fix: URL解析自动补全协议头

-uf 文件中 192.168.1.1:8080 自动转为 http://192.168.1.1:8080

* fix: 修复-u/-uf模式下URLs丢失导致0目标扫描的问题

Parse阶段将URLs设置到全局状态,但Initialize随后创建新状态
并覆盖了全局状态,导致URLs数据丢失。现在在创建新状态前
先保存并迁移Parse阶段设置的URLs和HostPorts数据。

* fix: 智能检测HTTP/HTTPS协议并优化URL显示

- 修复-u/-uf模式URLs丢失导致0目标扫描问题
- detectProtocol改为主动TLS握手检测,不依赖服务名
- WebTitle输出显示完整协议(http/https)
- 隐藏标准端口(80/443)使输出更简洁

* refactor: 精简parsers包,统一配置构建入口

- 删除冗余的中间层(XXXInput、XXXParser类)
- 新增 config_builder.go 统一配置构建
- parsers包从3000+行精简至~540行
- 保留核心函数:ParseIP、ParsePort、文件读取、凭据解析

* test: 扩展parsers单元测试覆盖边缘情况

- 新增内网简写解析测试(192/172/10)
- 新增完整IP范围和无效CIDR测试
- 新增Windows行尾(CRLF)处理测试
- 新增凭据和哈希文件解析测试
- 新增端口解析边缘情况测试
- 测试覆盖率达到94.2%

* refactor: 优化控制台输出格式

- 去掉时间戳,保留[*][+]前缀
- Web输出合并WebTitle和WebFinger为一行
- 有指纹显示绿色[+],无指纹显示白色[*]
- 格式: code:xxx len:xxx title:xxx server:xxx [指纹]
- 服务探测格式: [Product:xxx ||Version:xxx] Banner:(xxx)
- 字段对齐,输出更清爽

* feat: 添加凭据测试未发现弱密码的提示

- credential_tester.go: 失败时设置 Type=ResultTypeCredential
- scanner.go: 根据结果类型在 error 级别输出'未发现弱密码'提示
- 新增 i18n 翻译 brute_no_weak_pass

使用 -log all 或 -log error 可看到此提示

* refactor(logging): 重构日志级别为层级过滤设计

- LogLevel 从 string 改为 int 类型,支持层级比较
- 层级设计:Debug(0) < Base(1) < Info(2) < Success(3) < Vuln(4) < Error(5)
- 设置一个级别后,显示该级别及以上的日志
- Error 级别始终显示,不会被配置过滤掉
- 保留向后兼容别名(LevelAll, LevelInfoSuccess 等)
- 更新测试以匹配新的层级过滤行为

* style(logging): Error级别日志改为黄色显示

* style(findnet): NetInfo输出改为每行一个IP

* refactor(ms17010): 优化错误提示,明确指出SMBv1不支持等情况

* fix(credential): 修复凭据测试结果不一致的问题

问题原因:
1. 未知错误类型不重试,导致服务端限流时跳过正确密码
2. SSH 错误分类不够准确,某些临时错误未被识别

修复内容:
1. 未知错误改为可重试(可能是临时问题)
2. 增加 SSH 特有的网络错误识别(handshake failed, disconnect 等)

* fix(portfinger): 修复SMB2服务指纹识别和NetInfo输出问题

- 添加SMB2ProgNeg探针支持现代Windows的SMB2协议
- 修复Go regexp对高位字节的UTF-8兼容问题,使用Latin-1转换
- 修复探针失败后连接重建逻辑
- 修复vendor_product字段名不匹配问题
- 修复NetInfo多行输出被其他日志打断的问题

* fix(config): 从默认端口移除9100,避免触发打印机打印 (#517)

* feat(proxy): 增强代理端口扫描的深度验证机制

- 新增4阶段深度验证:Banner读取→探测发送→响应等待→最终判定
- 新增SOCKS5错误码和代理错误文本检测
- 优化ProbeProxyBehavior探测逻辑,发送数据验证连接可达性
- 解决透明代理/全回显代理导致的假阳性问题

* fix(proxy): 修复代理深度验证的若干问题

- detector.go: 修复 AutoConfigureProxy 覆盖探测结果的问题
  只有未探测过时才设置默认 proxyReliable 值

- port_scan.go: 改进深度验证机制
  - 使用带 Host header 的 HTTP GET 请求替代 OPTIONS
  - 延长响应等待超时至 2s 以适配慢速服务器
  - 正确重置连接 deadline 避免影响后续操作

* refactor: 统一 common 包文件命名风格

Flag.go -> flag.go

* refactor(proxy): 删除自定义 contains() 函数,改用标准库

- 用 strings.Contains() 替代手写的 contains()
- 删除过时的注释

* fix(parsers): 修复带横杠域名被误识别为IP范围的问题

如 111-555.sss.com 这类域名因包含 - 被错误解析为 IP 范围,
添加 looksLikeIPRange() 检查,只有 - 前是有效 IP 才走范围解析

* fix(proxy): 修复代理模式下服务识别错误和端口漏扫问题

- port_scan.go: 验证通过后重建干净连接,避免HTTP GET探测污染服务识别
- port_scan.go: 优化验证策略,用轻量CRLF探测替代HTTP GET,超时从2.2s降至0.6s
- manager.go: 修正ProbeProxyBehavior判断逻辑,超时应视为代理正常转发

* fix(pool): 移除线程池预分配,优化大规模扫描内存占用

WithPreAlloc(true) 会预先创建所有 worker goroutine,
在大规模扫描(如 25域名×65535端口)时可能导致内存问题

* refactor(logging): 统一日志前缀,删除废弃的 LogBase

- 删除 LogBase 函数,所有调用迁移到 LogInfo/LogError
- 新增 PrefixDebug ([.]) 前缀,所有日志级别现在都有前缀
- 修复日志输出缩进不一致的问题
- 删除未使用的 PrefixDefault 常量

* perf(icmp): 实现自适应等待算法优化存活检测性能

- 新增 waitAdaptive 函数,监控响应增量实现智能提前结束
- 算法保守原则:最小等待1s + 连续500ms无新响应才提前结束
- 添加100ms检查间隔避免CPU空转
- 保留原有最大等待时间(3s/6s)作为兜底
- 添加完整单元测试覆盖各种场景

优化效果:
- 全部响应:~100ms (原3s)
- 无响应:~1s (原3s)
- 部分响应后稳定:~1.5s (原3s)

* perf(scan): 实现启发式优化提升扫描体验

1. 端口优先级排序:高价值端口(80,443,22,3389等)优先扫描
   - 用户能更快看到有意义的结果
   - 不影响端口喷洒策略

2. TCP 补充探测:ICMP 响应率<10%时自动启用
   - 对未响应主机用 TCP 80/443/22/445 补充探测
   - 解决防火墙过滤 ICMP 导致漏检的问题

* refactor(grdp): 精简RDP库,删除认证检测不需要的代码

- 删除 VNC 协议支持 (protocol/rfb, client/rfb.go)
- 删除完整客户端框架 (client/)
- 删除 RemoteApp 等插件 (plugin/)
- 删除 RLE 图形解压 (core/rle.go)
- 删除绘图指令处理 (pdu/orders.go, pdu/gdi.go)
- 精简 screen.go,移除截图和完整会话功能
- 移除未使用的 RGB 转换函数

grdp 代码从 13,044 行精简至 7,581 行,削减 42%

* refactor(common): 删除死代码,优化代码风格

- 删除未使用的 joinStrings/joinInts 函数
- 删除未使用的 memStats 字段和 getMemoryInfo 方法
- 简化 parsePasswords 中的循环为 append(...) 形式

* refactor(services): 统一数据库插件的DBWrapper

4个数据库插件(MySQL、PostgreSQL、MSSQL、Oracle)都有相同的sql.DB包装代码,
合并为通用的SQLDBWrapper,减少重复。

* refactor(core,grdp): 删除未使用的死代码

- 移除 BaseScanStrategy.LogPluginInfoWithPort 方法(无调用者)
- 移除 mcs.go 中被注释的旧 connect 函数实现

* refactor: 删除 deadcode 检测出的未使用函数

- proxy/detector.go: 删除 IsSOCKS5Standard, IsProxyInitialized
- findnet.go: 删除 NetworkInfo.OneLine, TreeFormat 方法
- port_scan.go: 删除 estimateScanTime 函数
- web_scanner.go: 删除 GetFingerprints 函数
- 清理相关测试代码

* refactor: 删除更多未使用的死代码

- parse.go: 删除 RemoveDuplicate 函数及其测试
- parsers.go: 删除 excludeHosts, removeDuplicates 别名函数
- 更新测试使用真正的函数名

* fix(test): 修复 TestParseIP_InvalidIPRange 测试用例

- 删除不合理的测试用例(无效IP被当作普通主机名处理是设计行为)
- 修复测试逻辑,只在真正通过时输出"正确"

* fix(scan): 移除域名预解析,保留原始域名进行扫描

域名预解析会将域名转换为IP,导致虚拟主机场景下HTTP访问失败
(Host头变成IP而非域名,无法正确路由)

* fix(scan): 修复 -hf 参数无法单独使用的问题

* fix(proxy): 修复透明代理环境下 SOCKS5 代理全端口误报问题

问题:在透明代理(TUN模式)环境下使用 SOCKS5 代理扫描时,
会出现全端口开放的误报,因为代理可靠性检测被透明代理污染。

修复方案(参考 fscanx):
1. 将探针从 CRLF 改为 HTTP GET,更有效检测真实连接状态
2. 删除 "uncertain" 状态,无响应一律判定为端口关闭
3. 调整超时时间以适应代理链路延迟

Fixes #524

* feat(telnet): 新增 telnetd RCE 命令执行验证,修复未授权访问日志级别

* fix: 修复 i18n.Tr vet 报错、Unicode 测试用例,移除过期域名

- 移除 i18n.Tr 中错误的 fmt.Sprintf fallback,消除 go vet 误报
- 修复 match_engine_test Unicode 测试用例与 Latin-1 转换逻辑不匹配
- README 移除过期的 fscan.club 域名
- 添加 .gitattributes 统一换行符为 LF

* refactor: 统一控制台输出风格,使用统一的日志函数

手动合并 PR #558 的改动,适配重构后的代码路径

* fix(ci): 修复版本注入和CI触发配置

- goreleaser ldflags 指向正确的包路径 common.version/commit/date
- version 改为 var 支持 ldflags 注入,banner 显示 commit 和构建日期
- test-build 触发分支增加 dev-* 通配

* fix(ci): 修复 Windows 产物 .exe.exe 双后缀问题

* feat(ci): 扩展构建架构支持 MIPS/ARM/FreeBSD/Solaris
2026-04-25 17:39:16 +08:00

641 lines
17 KiB
Go

package fingerprint
import (
"testing"
)
/*
enhanced_test.go - Web指纹匹配引擎测试
测试重点:
1. matchWords - 关键词匹配逻辑(AND/OR条件、大小写)
2. matchFavicon - favicon hash匹配
3. CalculateFaviconHash - hash计算一致性
不测试:
- MatchEnhancedFingerprints - 依赖嵌入的JSON数据和全局状态
- matchRegex - 依赖全局regexCache,需要集成测试
*/
// =============================================================================
// matchWords 关键词匹配测试
// =============================================================================
// 创建测试用的matcher结构
func createMatcher(matcherType string, words, regex, hash []string, part, condition string, caseInsensitive bool) struct {
Type string `json:"type"`
Words []string `json:"words"`
Regex []string `json:"regex"`
Hash []string `json:"hash"`
Part string `json:"part"`
CaseInsensitive bool `json:"case-insensitive"`
Condition string `json:"condition"`
} {
return struct {
Type string `json:"type"`
Words []string `json:"words"`
Regex []string `json:"regex"`
Hash []string `json:"hash"`
Part string `json:"part"`
CaseInsensitive bool `json:"case-insensitive"`
Condition string `json:"condition"`
}{
Type: matcherType,
Words: words,
Regex: regex,
Hash: hash,
Part: part,
CaseInsensitive: caseInsensitive,
Condition: condition,
}
}
// TestMatchWords_ORCondition 测试OR条件匹配
func TestMatchWords_ORCondition(t *testing.T) {
tests := []struct {
name string
words []string
body string
expected bool
}{
{
name: "匹配第一个词",
words: []string{"nginx", "apache", "iis"},
body: "Server: nginx/1.18.0",
expected: true,
},
{
name: "匹配中间词",
words: []string{"nginx", "Apache", "iis"},
body: "Apache/2.4.41",
expected: true,
},
{
name: "匹配最后词",
words: []string{"nginx", "apache", "IIS"},
body: "Microsoft-IIS/10.0",
expected: true,
},
{
name: "无匹配",
words: []string{"nginx", "apache", "iis"},
body: "lighttpd/1.4.55",
expected: false,
},
{
name: "空body",
words: []string{"nginx"},
body: "",
expected: false,
},
{
name: "空words",
words: []string{},
body: "nginx",
expected: false,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
matcher := createMatcher("word", tt.words, nil, nil, "body", "", false)
result := matchWords(matcher, tt.body, "")
if result != tt.expected {
t.Errorf("matchWords() = %v, 期望 %v", result, tt.expected)
}
})
}
}
// TestMatchWords_ANDCondition 测试AND条件匹配
func TestMatchWords_ANDCondition(t *testing.T) {
tests := []struct {
name string
words []string
body string
expected bool
}{
{
name: "全部匹配",
words: []string{"WordPress", "wp-content", "wp-includes"},
body: "<html>WordPress site with wp-content and wp-includes</html>",
expected: true,
},
{
name: "部分匹配",
words: []string{"WordPress", "wp-content", "wp-includes"},
body: "WordPress site with wp-content",
expected: false,
},
{
name: "无匹配",
words: []string{"WordPress", "wp-content"},
body: "Joomla CMS",
expected: false,
},
{
name: "单词全匹配",
words: []string{"nginx"},
body: "nginx/1.18.0",
expected: true,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
matcher := createMatcher("word", tt.words, nil, nil, "body", "and", false)
result := matchWords(matcher, tt.body, "")
if result != tt.expected {
t.Errorf("matchWords(AND) = %v, 期望 %v", result, tt.expected)
}
})
}
}
// TestMatchWords_CaseInsensitive 测试大小写不敏感匹配
func TestMatchWords_CaseInsensitive(t *testing.T) {
tests := []struct {
name string
words []string
body string
caseInsensitive bool
expected bool
}{
{
name: "大小写敏感-精确匹配",
words: []string{"WordPress"},
body: "WordPress",
caseInsensitive: false,
expected: true,
},
{
name: "大小写敏感-不匹配",
words: []string{"WordPress"},
body: "wordpress",
caseInsensitive: false,
expected: false,
},
{
name: "大小写不敏感-小写匹配大写",
words: []string{"wordpress"},
body: "WORDPRESS",
caseInsensitive: true,
expected: true,
},
{
name: "大小写不敏感-混合大小写",
words: []string{"WoRdPrEsS"},
body: "wordpress site",
caseInsensitive: true,
expected: true,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
matcher := createMatcher("word", tt.words, nil, nil, "body", "", tt.caseInsensitive)
result := matchWords(matcher, tt.body, "")
if result != tt.expected {
t.Errorf("matchWords(caseInsensitive=%v) = %v, 期望 %v",
tt.caseInsensitive, result, tt.expected)
}
})
}
}
// TestMatchWords_HeaderPart 测试header部分匹配
func TestMatchWords_HeaderPart(t *testing.T) {
body := "<html>Body content</html>"
headers := "Server: nginx\r\nX-Powered-By: PHP/7.4"
tests := []struct {
name string
words []string
part string
expected bool
}{
{
name: "匹配header",
words: []string{"nginx"},
part: "header",
expected: true,
},
{
name: "header中不存在",
words: []string{"apache"},
part: "header",
expected: false,
},
{
name: "匹配body",
words: []string{"Body content"},
part: "body",
expected: true,
},
{
name: "body中不存在header内容",
words: []string{"nginx"},
part: "body",
expected: false,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
matcher := createMatcher("word", tt.words, nil, nil, tt.part, "", false)
result := matchWords(matcher, body, headers)
if result != tt.expected {
t.Errorf("matchWords(part=%s) = %v, 期望 %v",
tt.part, result, tt.expected)
}
})
}
}
// =============================================================================
// matchFavicon 测试
// =============================================================================
// TestMatchFavicon_Basic 测试favicon hash匹配
func TestMatchFavicon_Basic(t *testing.T) {
tests := []struct {
name string
hashes []string
favicon FaviconHashes
expected bool
}{
{
name: "mmh3匹配",
hashes: []string{"1386054408", "def456"},
favicon: FaviconHashes{MMH3: "1386054408", MD5: "abc"},
expected: true,
},
{
name: "MD5匹配",
hashes: []string{"abc123", "e2e2ba13339c2fea220f8b4fa6c32c0d"},
favicon: FaviconHashes{MMH3: "123", MD5: "e2e2ba13339c2fea220f8b4fa6c32c0d"},
expected: true,
},
{
name: "无匹配",
hashes: []string{"abc123", "def456"},
favicon: FaviconHashes{MMH3: "xyz789", MD5: "111"},
expected: false,
},
{
name: "空favicon",
hashes: []string{"abc123"},
favicon: FaviconHashes{},
expected: false,
},
{
name: "空hashes",
hashes: []string{},
favicon: FaviconHashes{MMH3: "abc123", MD5: "def"},
expected: false,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
matcher := createMatcher("favicon", nil, nil, tt.hashes, "", "", false)
result := matchFavicon(matcher, tt.favicon)
if result != tt.expected {
t.Errorf("matchFavicon() = %v, 期望 %v", result, tt.expected)
}
})
}
}
// =============================================================================
// CalculateFaviconHashes 测试
// =============================================================================
// TestCalculateFaviconHashes_Consistency 测试hash计算一致性
func TestCalculateFaviconHashes_Consistency(t *testing.T) {
data := []byte("test favicon data")
hash1 := CalculateFaviconHashes(data)
hash2 := CalculateFaviconHashes(data)
if hash1.MMH3 != hash2.MMH3 {
t.Errorf("相同数据应产生相同mmh3 hash: %s vs %s", hash1.MMH3, hash2.MMH3)
}
if hash1.MD5 != hash2.MD5 {
t.Errorf("相同数据应产生相同MD5 hash: %s vs %s", hash1.MD5, hash2.MD5)
}
}
// TestCalculateFaviconHashes_Different 测试不同数据产生不同hash
func TestCalculateFaviconHashes_Different(t *testing.T) {
data1 := []byte("favicon data 1")
data2 := []byte("favicon data 2")
hash1 := CalculateFaviconHashes(data1)
hash2 := CalculateFaviconHashes(data2)
if hash1.MMH3 == hash2.MMH3 {
t.Error("不同数据应产生不同mmh3 hash")
}
if hash1.MD5 == hash2.MD5 {
t.Error("不同数据应产生不同MD5 hash")
}
}
// TestCalculateFaviconHashes_Empty 测试空数据
func TestCalculateFaviconHashes_Empty(t *testing.T) {
hash := CalculateFaviconHashes([]byte{})
if hash.MMH3 != "" || hash.MD5 != "" {
t.Errorf("空数据应返回空FaviconHashes,实际: mmh3=%s, md5=%s", hash.MMH3, hash.MD5)
}
}
// TestCalculateFaviconHashes_Nil 测试nil数据
func TestCalculateFaviconHashes_Nil(t *testing.T) {
hash := CalculateFaviconHashes(nil)
if hash.MMH3 != "" || hash.MD5 != "" {
t.Errorf("nil数据应返回空FaviconHashes,实际: mmh3=%s, md5=%s", hash.MMH3, hash.MD5)
}
}
// TestCalculateFaviconHashes_Format 测试hash格式
func TestCalculateFaviconHashes_Format(t *testing.T) {
data := []byte("test data")
hash := CalculateFaviconHashes(data)
// mmh3 应该是有符号整数格式(可能是负数)
if hash.MMH3 == "" {
t.Error("mmh3 hash不应为空")
}
// MD5产生32字符的十六进制字符串
if len(hash.MD5) != 32 {
t.Errorf("MD5 hash长度应为32,实际: %d", len(hash.MD5))
}
// 验证MD5是有效的十六进制
for _, c := range hash.MD5 {
if (c < '0' || c > '9') && (c < 'a' || c > 'f') {
t.Errorf("MD5 hash包含无效字符: %c", c)
}
}
}
// TestMMH3_KnownValue 测试mmh3已知值(验证算法正确性)
func TestMMH3_KnownValue(t *testing.T) {
// 使用简单的测试字符串验证mmh3算法
// mmh3("hello") with seed=0 应该产生一个固定值
result := mmh3Hash32([]byte("hello"))
// mmh3("hello", seed=0) = 613153351 (根据标准实现)
expected := int32(613153351)
if result != expected {
t.Errorf("mmh3('hello') = %d, 期望 %d", result, expected)
}
}
// =============================================================================
// matchMatcher 分发测试
// =============================================================================
// TestMatchMatcher_TypeDispatch 测试类型分发
func TestMatchMatcher_TypeDispatch(t *testing.T) {
// 注意:matchRegex需要全局regexCache,这里只测试word和favicon
tests := []struct {
name string
matcherType string
expected bool
}{
{
name: "word类型",
matcherType: "word",
expected: true,
},
{
name: "favicon类型",
matcherType: "favicon",
expected: true,
},
{
name: "未知类型",
matcherType: "unknown",
expected: false,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
var matcher struct {
Type string `json:"type"`
Words []string `json:"words"`
Regex []string `json:"regex"`
Hash []string `json:"hash"`
Part string `json:"part"`
CaseInsensitive bool `json:"case-insensitive"`
Condition string `json:"condition"`
}
switch tt.matcherType {
case "word":
matcher = createMatcher("word", []string{"nginx"}, nil, nil, "body", "", false)
case "favicon":
matcher = createMatcher("favicon", nil, nil, []string{"abc123"}, "", "", false)
default:
matcher = createMatcher(tt.matcherType, nil, nil, nil, "", "", false)
}
result := matchMatcher(matcher, "nginx server", "Server: nginx", FaviconHashes{MMH3: "abc123", MD5: "def456"}, nil)
if result != tt.expected {
t.Errorf("matchMatcher(type=%s) = %v, 期望 %v",
tt.matcherType, result, tt.expected)
}
})
}
}
// =============================================================================
// 边界情况测试
// =============================================================================
// TestMatchWords_SpecialCharacters 测试特殊字符
func TestMatchWords_SpecialCharacters(t *testing.T) {
tests := []struct {
name string
words []string
body string
expected bool
}{
{
name: "包含点号",
words: []string{"nginx/1.18.0"},
body: "Server: nginx/1.18.0",
expected: true,
},
{
name: "包含括号",
words: []string{"(Ubuntu)"},
body: "Apache/2.4.41 (Ubuntu)",
expected: true,
},
{
name: "包含中文",
words: []string{"欢迎"},
body: "<title>欢迎访问</title>",
expected: true,
},
{
name: "包含换行符",
words: []string{"Content-Type"},
body: "Header:\r\nContent-Type: text/html",
expected: true,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
matcher := createMatcher("word", tt.words, nil, nil, "body", "", false)
result := matchWords(matcher, tt.body, "")
if result != tt.expected {
t.Errorf("matchWords() = %v, 期望 %v", result, tt.expected)
}
})
}
}
// TestMatchWords_LargeBody 测试大body
func TestMatchWords_LargeBody(t *testing.T) {
// 构造100KB的body
largeBody := make([]byte, 100*1024)
for i := range largeBody {
largeBody[i] = 'x'
}
// 在中间插入关键词
copy(largeBody[50*1024:], []byte("WordPress"))
matcher := createMatcher("word", []string{"WordPress"}, nil, nil, "body", "", false)
result := matchWords(matcher, string(largeBody), "")
if !result {
t.Error("大body中的关键词应被匹配")
}
}
// =============================================================================
// 性能基准测试
// =============================================================================
// =============================================================================
// 版本提取测试
// =============================================================================
// TestExtractVersions_ServerHeaders 测试从 Server 头提取版本
func TestExtractVersions_ServerHeaders(t *testing.T) {
tests := []struct {
name string
headers string
expected map[string]string // name -> version
}{
{
name: "nginx版本",
headers: "Server: nginx/1.18.0",
expected: map[string]string{"nginx": "1.18.0"},
},
{
name: "Apache版本",
headers: "Server: Apache/2.4.41 (Ubuntu)",
expected: map[string]string{"apache": "2.4.41"},
},
{
name: "多个版本",
headers: "Server: nginx/1.18.0\nX-Powered-By: PHP/7.4.3",
expected: map[string]string{"nginx": "1.18.0", "php": "7.4.3"},
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
results := ExtractVersions("", tt.headers)
for _, v := range results {
if expected, ok := tt.expected[v.Name]; ok {
if v.Version != expected {
t.Errorf("%s 版本不匹配: got %s, want %s", v.Name, v.Version, expected)
}
}
}
})
}
}
// TestExtractVersions_BodyContent 测试从 body 提取版本
func TestExtractVersions_BodyContent(t *testing.T) {
body := `
<!DOCTYPE html>
<html>
<head>
<meta name="generator" content="WordPress 6.0" />
<script src="/js/jquery-3.6.0.min.js"></script>
</head>
<body>Powered by Tomcat/9.0.41</body>
</html>`
results := ExtractVersions(body, "")
// 检查是否提取到预期的版本
found := make(map[string]bool)
for _, v := range results {
found[v.Name] = true
t.Logf("提取到: %s %s", v.Name, v.Version)
}
// WordPress 可能无法提取(因为正则需要调整),但 jQuery 和 Tomcat 应该可以
if !found["jquery"] && !found["tomcat"] {
t.Error("应至少提取到 jquery 或 tomcat 版本")
}
}
// TestExtractVersions_Empty 测试空输入
func TestExtractVersions_Empty(t *testing.T) {
results := ExtractVersions("", "")
if len(results) != 0 {
t.Errorf("空输入应返回空结果,实际: %d", len(results))
}
}
// BenchmarkExtractVersions 基准测试:版本提取性能
func BenchmarkExtractVersions(b *testing.B) {
headers := "Server: nginx/1.18.0\nX-Powered-By: PHP/8.0.3\n"
body := `<meta name="generator" content="WordPress 6.0" />`
b.ResetTimer()
for i := 0; i < b.N; i++ {
_ = ExtractVersions(body, headers)
}
}
// BenchmarkMatchEnhancedFingerprints 基准测试:并发指纹匹配
func BenchmarkMatchEnhancedFingerprints(b *testing.B) {
// 模拟真实的 HTTP 响应
body := []byte(`<!DOCTYPE html>
<html>
<head><title>WordPress Site</title></head>
<body>
<meta name="generator" content="WordPress 6.0" />
<link rel="stylesheet" href="/wp-content/themes/flavor/style.css" />
Powered by nginx/1.18.0
</body>
</html>`)
headers := "Server: nginx/1.18.0\nX-Powered-By: PHP/8.0\n"
favicon := FaviconHashes{MMH3: "1386054408", MD5: "abc123"}
// 预热:确保指纹库已加载
_ = MatchEnhancedFingerprints(body, headers, favicon)
b.ResetTimer()
for i := 0; i < b.N; i++ {
_ = MatchEnhancedFingerprints(body, headers, favicon)
}
}