Conversation
BPF_PROG_LOAD handed the raw user instruction stream straight to the relocation pass, which treated every eight-byte slot as a complete instruction. A stream whose last slot was `BPF_LD | BPF_IMM | BPF_DW` therefore read the slot past the end of the stream and panicked inside the kernel; the panic was swallowed by the syscall unwind handler, so userspace observed ENOSYS instead of the EINVAL Linux 6.6 returns from `resolve_pseudo_ldimm64()`. The same missing width model let a plain 64-bit immediate, and an unknown pseudo-register class, be advanced by a single slot and silently published as a valid program. Model the instruction layout explicitly instead of patching the bound: * Add `bpf/prog/instructions.rs`, which owns the slot width rules: two-slot `LD_DW_IMM` structure checks, the `LDX` reserved-field rule, the opcode table taken from Linux, `ldimm64` classification and the in-place immediate write-back. The rules live in one place and are unit tested without building a whole program. * Rewrite relocation as a single pass that validates and relocates at the same time, so a malformed stream fails before a program fd exists. The errno follows Linux for each case: EINVAL for a truncated or malformed pair, EPROTO for `BPF_PSEUDO_MAP_IDX*` without an fd array, ENOTSUPP for a map type without a direct value address. * Replace `BpfProg::raw_file_ptr` (`Arc::into_raw`/`Arc::from_raw` balanced by `Drop`) with `Vec<Arc<BpfMap>>`. The program keeps one reference per distinct map, deduplicated and bounded by MAX_USED_MAPS, so a failure in the middle of relocation cannot leak a map and a closed descriptor cannot leave a dangling instruction immediate. * Converge the map capability on `direct_value_ptr(offset)`: only a single-entry array map answers, with the offset bounded by the requested value size, which mirrors `array_map_direct_value_addr()`. Remove the `first_value_ptr()` implementations that returned the data address for map types Linux refuses. * Size the array map buffer in `usize` with a checked multiply and a fallible allocation, and track the requested `value_size` separately from the rounded element size. * Let the nine raw helpers borrow the map (`&*map`) instead of rebuilding an `Arc` from a pointer and relying on the balanced `into_raw` call. * Compute the program tag before relocation so it covers the program as submitted. Add `normal/bpf_prog_load` to dunitest (14 cases) and to the whitelist. On the host and on the DragonOS guest the same cases pass, the original probe now returns EINVAL without panicking, and the BPF regression suites (`bpf_syscall_abi`, `cgroup_device_bpf`, `perf_bpf_mmap`) still pass. The cgroup device verifier and its constraints are untouched. Signed-off-by: longjin <longjin@dragonos.org>
Member
Author
|
@codex review |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
背景
bpf(BPF_PROG_LOAD)把用户提交的指令流直接交给重定位过程,而重定位把每个 8 字节槽都当成一条完整指令。于是:BPF_LD | BPF_IMM | BPF_DW(双槽 64 位立即数加载)结尾时,会读取末尾之后不存在的槽。这是用户可控输入触发的内核越界,panic 被系统调用展开路径吞掉后,用户态看到ENOSYS,而 Linux 6.6 在resolve_pseudo_ldimm64()里以EINVAL拒绝不完整指令。此外,重定位路径上的 map 生命周期是手工用
Arc::into_raw/Arc::from_raw平衡的:中途失败会泄漏已取得的引用;BPF_PSEUDO_MAP_VALUE不持有 map 引用且不校验偏移;helper 用from_raw/into_raw假装借用。这些是同一个"指令流布局准入 + 双槽重定位 + map 引用生命周期"功能的缺失,因此本次整体实现,而不是只加一个下标判断。修复内容
kernel/src/bpf/prog/instructions.rs,作为指令布局规则的唯一归属:双槽LD_DW_IMM结构校验、LDX保留字段规则、来自 Linux 的 opcode 表、ldimm64分类与就地立即数写回;规则可脱离完整程序单独单测。EINVAL,无 fd_array 的BPF_PSEUDO_MAP_IDX*为EPROTO,没有直接取值地址的 map 类型为ENOTSUPP)。BpfProg用Vec<Arc<BpfMap>>取代raw_file_ptr:每个去重后的 map 持有一个引用并受MAX_USED_MAPS约束,重定位中途失败不再泄漏,关闭 map fd 也不会让指令里的地址悬垂。direct_value_ptr(offset):只有单元素 array map 应答,offset以用户请求的value_size为上界,对应 Linuxarray_map_direct_value_addr();删除对 Linux 不支持的 map 类型返回数据首地址的first_value_ptr()。usize计算、checked_mul与可失败分配,并单独记录请求的value_size(与向上取整后的元素大小区分)。&*map),不再重建Arc。验证
make kernel通过,改动文件无新增 warning;cargo fmt --all -- --check通过。normal/bpf_prog_load14/14 通过。ENOSYS变为fd=-1 errno=22,退出码 0;bpf_prog_load_test14/14;bpf_syscall_abi_test10/10、cgroup_device_bpf_test5/5、perf_bpf_mmap_test2/2。normal/bpf_prog_load已加入whitelist.txt,随默认 dunitest 批次进入 CI;无 cap 时退化为 skip。明确不在本次范围
key_size == 0触发assert_ne!panic、PERF_EVENT_ARRAY 越界 key 切片 panic、QUEUE/STACK 无界不可失败分配)与本次改动路径无因果关系,另列专项处理。cgroup device 走独立受限验证器,其约束在本次改动中保持不变。