Skip to content

fix(benchmarks): use nullMarkers in csvOptions and add missing request counter schema fields - #1062

Merged
zhixiangli merged 2 commits into
fsspec:mainfrom
yuxin00j:fix-benchmarks-schema-cat-ranges
Sep 18, 2026
Merged

zhixiangli merged 2 commits into
fsspec:mainfrom
yuxin00j:fix-benchmarks-schema-cat-ranges

Conversation

@yuxin00j

@yuxin00j yuxin00j commented Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  1. Change "nullMarker": "N/A" to "nullMarkers": ["N/A", ""] in cloudbuild/benchmarks/benchmarks_schema.json (matching macrobenchmarks_schema.json and subsystembenchmarks_schema.json) so BigQuery treats empty strings from missing columns in historical CSVs as NULL, preserving numeric types (INTEGER / FLOAT).
  2. Add missing schema fields introduced in Adds cat file benchmarks #1052 (concurrency, requests_total, requests_per_op, requests_download, requests_object_get, requests_list, requests_other).

Problem

  1. When BigQuery scans historical CSV files generated prior to feat(perf): add microbenchmark suite for cat_ranges #1035 under "sourceColumnMatch": "NAME", missing columns evaluate to empty strings (""). With "nullMarker": "N/A", BigQuery did not treat "" as NULL, failing ingestion when parsing "" into INTEGER columns (batch_size, max_gap, num_ranges, total_bytes). Using "nullMarkers": ["N/A", ""] resolves this while preserving INTEGER types.
  2. Recent microbenchmark runs after Adds cat file benchmarks #1052 emit request counter metrics (concurrency, requests_total, requests_per_op, requests_download, requests_object_get, requests_list, requests_other) in CSV headers, which caused BigQuery ingestion to fail with Unmatched source column found when scanning newer CSVs.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request changes the data types of several fields (batch_size, max_gap, num_ranges, and total_bytes) from INTEGER to STRING in the benchmarks schema. The reviewer suggests keeping these fields as INTEGER and instead configuring csvOptions to treat empty strings as NULL, which preserves schema quality and proper numeric types for querying.

Comment thread cloudbuild/benchmarks/benchmarks_schema.json Outdated
@codecov

codecov Bot commented Sep 14, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 90.21%. Comparing base (581fd26) to head (45dc6bb).

Additional details and impacted files
@@           Coverage Diff           @@
##             main    #1062   +/-   ##
=======================================
  Coverage   90.21%   90.21%           
=======================================
  Files          16       16           
  Lines        3749     3749           
=======================================
  Hits         3382     3382           
  Misses        367      367           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@yuxin00j
yuxin00j marked this pull request as draft September 14, 2026 05:14
@yuxin00j yuxin00j changed the title fix(benchmarks): change cat_ranges schema fields to STRING for BigQuery ingestion fix(benchmarks): use nullMarkers in csvOptions and add missing request counter schema fields Sep 14, 2026
@yuxin00j
yuxin00j requested a review from zhixiangli September 14, 2026 05:34
@yuxin00j
yuxin00j marked this pull request as ready for review September 14, 2026 05:34
@zhixiangli
zhixiangli merged commit 03e588e into fsspec:main Sep 18, 2026
11 of 12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants