Skip to content

fix(bpf): validate ldimm64 layout and hold map references - #2387

Open
fslongjin wants to merge 1 commit into
DragonOS-Community:masterfrom
fslongjin:codex/dkc005-bpf-ldimm64-validation
Open

fslongjin wants to merge 1 commit into
DragonOS-Community:masterfrom
fslongjin:codex/dkc005-bpf-ldimm64-validation

Conversation

@fslongjin

Copy link
Copy Markdown
Member

背景

bpf(BPF_PROG_LOAD) 把用户提交的指令流直接交给重定位过程,而重定位把每个 8 字节槽都当成一条完整指令。于是:

  • 指令流以 BPF_LD | BPF_IMM | BPF_DW(双槽 64 位立即数加载)结尾时,会读取末尾之后不存在的槽。这是用户可控输入触发的内核越界,panic 被系统调用展开路径吞掉后,用户态看到 ENOSYS,而 Linux 6.6 在 resolve_pseudo_ldimm64() 里以 EINVAL 拒绝不完整指令。
  • 同一个"缺少指令宽度模型"的问题还让普通 64 位立即数只推进一个槽(把立即数高 32 位当成指令解码),并让未实现的伪寄存器类别静默当作普通整数、以"加载成功"发布 fd。

此外,重定位路径上的 map 生命周期是手工用 Arc::into_raw/Arc::from_raw 平衡的:中途失败会泄漏已取得的引用;BPF_PSEUDO_MAP_VALUE 不持有 map 引用且不校验偏移;helper 用 from_raw/into_raw 假装借用。这些是同一个"指令流布局准入 + 双槽重定位 + map 引用生命周期"功能的缺失,因此本次整体实现,而不是只加一个下标判断。

修复内容

  • 新增 kernel/src/bpf/prog/instructions.rs,作为指令布局规则的唯一归属:双槽 LD_DW_IMM 结构校验、LDX 保留字段规则、来自 Linux 的 opcode 表、ldimm64 分类与就地立即数写回;规则可脱离完整程序单独单测。
  • 重定位改为单遍扫描:校验与重定位同时完成,畸形流在发布 fd 之前就失败。errno 逐项对齐 Linux(截断/畸形对 EINVAL,无 fd_array 的 BPF_PSEUDO_MAP_IDX* 为 EPROTO,没有直接取值地址的 map 类型为 ENOTSUPP)。
  • BpfProg 用 Vec<Arc<BpfMap>> 取代 raw_file_ptr:每个去重后的 map 持有一个引用并受 MAX_USED_MAPS 约束,重定位中途失败不再泄漏,关闭 map fd 也不会让指令里的地址悬垂。
  • map 能力收敛为 direct_value_ptr(offset):只有单元素 array map 应答,offset 以用户请求的 value_size 为上界,对应 Linux array_map_direct_value_addr();删除对 Linux 不支持的 map 类型返回数据首地址的 first_value_ptr()。
  • array map 缓冲区改用 usize 计算、checked_mul 与可失败分配,并单独记录请求的 value_size(与向上取整后的元素大小区分)。
  • 九个 raw helper 改为借用 map(&*map),不再重建 Arc。

验证

  • make kernel 通过,改动文件无新增 warning;cargo fmt --all -- --check 通过。
  • 宿主:新增套件 normal/bpf_prog_load 14/14 通过。
  • guest(本次改动最终编译出的内核,2 vCPU、快照盘、virtconsole socket):
    • 原 DKC-005 探针从 panic/ENOSYS 变为 fd=-1 errno=22,退出码 0;
    • bpf_prog_load_test 14/14;
    • 回归 bpf_syscall_abi_test 10/10、cgroup_device_bpf_test 5/5、perf_bpf_mmap_test 2/2。
  • normal/bpf_prog_load 已加入 whitelist.txt,随默认 dunitest 批次进入 CI;无 cap 时退化为 skip。
  • 已按架构/正确性、安全/边界、可维护性三个视角做过对抗性复核,并对评审意见逐条核实后再处置。

明确不在本次范围

  • 通用 eBPF 寄存器类型、内存访问与控制流(含 CFG)验证:属完整验证器,另列专项。
  • helper 实参类型校验:需要验证器提供类型信息才能闭环,另列专项;本次只把 helper 的借用前提写清楚。
  • 复核中确认的邻近既有缺陷(key_size == 0 触发 assert_ne! panic、PERF_EVENT_ARRAY 越界 key 切片 panic、QUEUE/STACK 无界不可失败分配)与本次改动路径无因果关系,另列专项处理。

cgroup device 走独立受限验证器,其约束在本次改动中保持不变。

BPF_PROG_LOAD handed the raw user instruction stream straight to the
relocation pass, which treated every eight-byte slot as a complete
instruction. A stream whose last slot was `BPF_LD | BPF_IMM | BPF_DW`
therefore read the slot past the end of the stream and panicked inside
the kernel; the panic was swallowed by the syscall unwind handler, so
userspace observed ENOSYS instead of the EINVAL Linux 6.6 returns from
`resolve_pseudo_ldimm64()`. The same missing width model let a plain
64-bit immediate, and an unknown pseudo-register class, be advanced by a
single slot and silently published as a valid program.

Model the instruction layout explicitly instead of patching the bound:

* Add `bpf/prog/instructions.rs`, which owns the slot width rules:
  two-slot `LD_DW_IMM` structure checks, the `LDX` reserved-field rule,
  the opcode table taken from Linux, `ldimm64` classification and the
  in-place immediate write-back. The rules live in one place and are
  unit tested without building a whole program.
* Rewrite relocation as a single pass that validates and relocates at the
  same time, so a malformed stream fails before a program fd exists. The
  errno follows Linux for each case: EINVAL for a truncated or malformed
  pair, EPROTO for `BPF_PSEUDO_MAP_IDX*` without an fd array, ENOTSUPP for
  a map type without a direct value address.
* Replace `BpfProg::raw_file_ptr` (`Arc::into_raw`/`Arc::from_raw`
  balanced by `Drop`) with `Vec<Arc<BpfMap>>`. The program keeps one
  reference per distinct map, deduplicated and bounded by MAX_USED_MAPS,
  so a failure in the middle of relocation cannot leak a map and a closed
  descriptor cannot leave a dangling instruction immediate.
* Converge the map capability on `direct_value_ptr(offset)`: only a
  single-entry array map answers, with the offset bounded by the
  requested value size, which mirrors `array_map_direct_value_addr()`.
  Remove the `first_value_ptr()` implementations that returned the data
  address for map types Linux refuses.
* Size the array map buffer in `usize` with a checked multiply and a
  fallible allocation, and track the requested `value_size` separately
  from the rounded element size.
* Let the nine raw helpers borrow the map (`&*map`) instead of rebuilding
  an `Arc` from a pointer and relying on the balanced `into_raw` call.
* Compute the program tag before relocation so it covers the program as
  submitted.

Add `normal/bpf_prog_load` to dunitest (14 cases) and to the whitelist.
On the host and on the DragonOS guest the same cases pass, the original
probe now returns EINVAL without panicking, and the BPF regression
suites (`bpf_syscall_abi`, `cgroup_device_bpf`, `perf_bpf_mmap`) still
pass. The cgroup device verifier and its constraints are untouched.

Signed-off-by: longjin <longjin@dragonos.org>
@fslongjin

Copy link
Copy Markdown
Member Author

@codex review

@github-actions github-actions Bot added the Bug fix A bug is fixed in this pull request label Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Bug fix A bug is fixed in this pull request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant