Skip to content

feat: .swd binary format + COS object storage integration - #1

Open
pescn wants to merge 386 commits into
mainfrom
feat/swd-binary-format
Open

pescn wants to merge 386 commits into
mainfrom
feat/swd-binary-format

Conversation

@pescn

@pescn pescn commented Feb 4, 2026

Copy link
Copy Markdown
Owner

Summary

  • Phase 1: 新增 .swd 定长二进制格式(128B Header + 24B 标量记录 / 80B 媒体索引记录),支持 O(1) 随机访问、HTTP Range GET 增量读取、崩溃截断恢复
  • Phase 2: 本地存储迁移 — ManifestManager、特性开关 (use_swd_format / use_cos_direct)、SwanLabKey 双写 JSONL + .swd、LocalRunCallback manifest 更新
  • Phase 3: COS 对象存储集成 — CosClient、CosAppender(position 校验 + 指数退避重试)、CloudPyCallback COS 全生命周期管理
  • Phase 4: Manifest 生命周期 — 定期 COS 同步 Timer、ConsoleLogAppender 4KB 缓冲追加上传
  • Phase 5: 向后兼容 — BackupToSwdConverter、JsonlToSwdConverter(支持双 key 格式)、SwdExporter (CSV/JSONL/Parquet)、swanlab sync .swd 文件检测与增量续传

新增文件

模块 文件 职责
格式定义 swanlab/data/swd/format.py 常量、struct 格式、NamedTuple
写入器 swanlab/data/swd/writer.py SwdWriter
读取器 swanlab/data/swd/reader.py SwdReader(文件 + bytes 双模式)
注册表 swanlab/data/swd/manifest.py ManifestManager(原子写入)
控制台 swanlab/data/swd/console.py ConsoleLogAppender
转换器 swanlab/data/swd/converter.py Backup/JSONL → SWD、SWD → CSV/JSONL/Parquet
格式文档 swanlab/data/swd/swd-format.md 格式规格说明
COS 客户端 swanlab/core_python/cos/client.py CosClient
COS 追加器 swanlab/core_python/cos/appender.py CosAppender
重试工具 swanlab/core_python/cos/retry.py 指数退避重试

修改文件

  • swanlab/swanlab_settings.py — 特性开关
  • swanlab/data/store.py — swd_dir、manifest_file、cos_config 属性
  • swanlab/data/sdk.py — 初始化时创建 swd 目录
  • swanlab/data/run/key.py — SwanLabKey 写入 .swd 数据
  • swanlab/data/run/exp.py — 创建 SwdWriter
  • swanlab/toolkit/models/metric.py — MetricInfo 增加 swd_bytes
  • swanlab/data/porter/__init__.py — trace_metric_swd()
  • swanlab/data/callbacker/cloud.py — COS 全生命周期
  • swanlab/data/callbacker/local.py — 双写 + manifest
  • swanlab/sync/__init__.py — .swd 检测与增量同步

特性开关

两个开关默认关闭,零风险合入:

use_swd_format: StrictBool = False   # 本地 .swd 存储
use_cos_direct: StrictBool = False   # COS 直接上传

Test plan

  • 98 个单元测试全部通过(format × 32、writer × 18、reader × 23、converter × 7、COS × 18)
  • 集成测试:swanlab.init() + swanlab.log() × 10000 + swanlab.finish() 验证 .swd 文件可读
  • swanlab sync 旧格式目录 → 验证 .swd 检测与 COS 续传
  • 前端 Range GET 增量读取验证

🤖 Generated with Claude Code

Zeyi-Lin and others added 30 commits March 24, 2025 11:04
* Fix: Change Dict to dict in _TYPE_HANDLERS type hints

* fix: allow dict type input for Object3D

The data can be passed in the form of dictionary.
For example: {"points": points_xyz, "boxes": "..."}.

* docs: add docs point clouds with boxes as input to Object3D

* docs: Add API reference link for Object3D
* feat: add discord webhook

* feat: support slack callback

* feat: add resp status_code
* fix(docs): en redirection

* fix: unused url

* fix(docs): unordered list
* feat: gpu time of memory

* test: for gpu mem time
* refactor: settings

* feat: enable conda and attach some other settings

* fix: iter settings

* fix: handle some bugs
* add cambricon npu detect
* fix: move cambricon from npu to mlu

* fix: opt cambricon codes

* fix
* feat: Add Molecule class for 3D molecule visualization

This commit introduces the Molecule class, which is used to visualize
3D molecules. The class can be initialized with PDB data or from an
RDKit Mol object. It includes methods for parsing the molecule data
into a buffer and for returning metadata for visualization.

* feat: Add molecule visualization test

This commit introduces a test case for visualizing molecules
using swanlab.data.modules.object3d.Molecule.

* chore: english comment.

* Add Molecule class to object3d module

This commit introduces the Molecule class for handling molecule
data from various formats using the RDKit library, and add it to
`__all__` in `__init__.py`.

* Feat: Add Molecule class methods for file inputs

Adds from_pdb_file, from_sdf_file, from_smiles and from_mol_file
methods to the Molecule class. These methods allow creating
Molecule instances from various file formats (PDB, SDF, Mol) and
SMILES strings, providing more flexibility in how molecule data is
loaded and displayed.

* Support molecule objects in Object3D

Enable creation of Object3D from rdkit.Chem.Mol objects and
add file handlers for common molecule file formats.

* test: add pdb file molecule test example.

* chore: update molecule test example.

* chore: change test example molecule file path.

* chore: add requirements rdkit for molecule

* chore: update code sytle for import.

* test: Add test for `swanlab.data.modules.object3d.Molecule`

This commit introduces a comprehensive test suite for the
`Molecule` class within the `swanlab.data.modules.object3d` module.
The tests cover various functionalities including:

- Creating `Molecule` objects from different file formats (PDB, SDF,
  Mol) and SMILES strings.
- Parsing `Molecule` objects.
- Verifying metadata such as caption and chart type.
- Ensures that `Object3D` correctly dispatches `Molecule` objects.

* chore: some misc

* chore: change chart type

* fix: test

---------

Co-authored-by: KAAANG <[email protected]>
* changelog evalscope

* changelog molecule

* support meta data type

* cambricon mlu

* hardware log

* fix typo

* hardware log en

* update example

* example lang

* add self-hosted link

* swanlab-architecture
* feat:add func of get uv requirements

* refactor: merge func of get uv  to get requirements

* chore:del uv settings

* test:add uv elapsed time test

* test:add try except
SAKURA-CAT and others added 30 commits December 5, 2025 17:53
* Add LOG_LEVEL env var for log level configuration

Introduced SWANLAB_LOG_LEVEL environment variable to control SwanLab's log output level. The log level now prioritizes the environment variable over the log_level argument, improving configurability.

* Refactor metrics upload logic and add batch uploader

Moved metrics upload batching logic from upload.py to a new batch.py module, introducing improved chunking and client state management. Updated related imports and test cases to use the new batch uploader. Enhanced logging in client/__init__.py and improved type handling in toolkit/logger.py.

* Improve logger level handling and add tests

Refactored logger initialization to support environment variable for log level and improved handling of invalid log levels. Added unit tests to verify environment variable log level and defaulting to 'info' for invalid levels.

* Refactor batch uploader to simplify request handling

Removed the client_state_guard decorator and integrated client state checks directly into the trace_metrics function. This streamlines the request logic and eliminates unnecessary indirection, making the code easier to follow and maintain.

* Refactor SwanLog tests for better isolation and flexibility

Updated test_log.py to use instance-level logger management, allowing tests to operate on different SwanLog instances and improving test isolation. Modified start_proxy to accept a logger parameter and return both log file and logger. Minor improvements in batch.py to handle empty data in trace_metrics and expanded type hints.

* Remove LOG_LEVEL from environment in reset_some_env

Ensures that the LOG_LEVEL environment variable is deleted if present when resetting environment variables in reset_some_env.

* Improve log level handling and test cleanup

Log level is now consistently set to lowercase in SwanLog. Test setup now resets swanlog to ensure proper cleanup, handling potential RuntimeError. Removed LOG_LEVEL env var deletion from reset_some_env for better environment management.

* Refactor uploader URL usage and update tests

Centralized the '/house/metrics' URL as HOUSE_URL in upload.py and updated all references. Improved logging in client/__init__.py and batch.py. Fixed and clarified test for trace_metrics behavior when client is pending.

* Fix handling of None data in trace_metrics

Updated trace_metrics to handle cases where data is None by defaulting to an empty list, preventing errors when calculating total_len.
* Initial plan

* Fix tutils check.py to handle missing swanboard package

Co-authored-by: SAKURA-CAT <[email protected]>

* Add comprehensive unit tests for TimeoutHTTPAdapter

Co-authored-by: SAKURA-CAT <[email protected]>

* Refactor tests to use HTTPAdapter directly instead of __bases__[0]

Co-authored-by: SAKURA-CAT <[email protected]>

* Remove duplicate imports and consolidate at top of file

Co-authored-by: SAKURA-CAT <[email protected]>

* Remove unnecessary @responses.activate decorators and move urllib3 import to top

Co-authored-by: SAKURA-CAT <[email protected]>

* Fix import order to follow PEP8 conventions

Co-authored-by: SAKURA-CAT <[email protected]>

* Fix swanboard version parsing in check.py

Improves the extraction of the swanboard version from pip output by handling whitespace and line endings more robustly.

---------

Co-authored-by: copilot-swe-agent[bot] <[email protected]>
Co-authored-by: SAKURA-CAT <[email protected]>
Co-authored-by: Kang Li <[email protected]>
* Increase max description length to 1024 characters

Updated the check_desc_format function to allow descriptions up to 1024 characters instead of 255, accommodating longer input as needed.

* Increase job_type max length to 256

Updated the maximum allowed length for job_type from 255 to 256 in check_job_type_format to accommodate longer job type strings.
* Remove VSCode config and documentation files

Deleted all files in the .vscode directory and several documentation markdown files. This cleans up editor-specific settings and outdated or unnecessary documentation from the repository.

* Add linter directives and project dictionary

Added a custom dictionary for the word 'noctx' in the project settings. Suppressed 'noctx' and 'revive' linter warnings in core.go and api.go, respectively, to address linter feedback and document intentional code decisions.

* Add comments for future context and linter exceptions

Added explanatory comments regarding future use of context in core.go and clarified nolint:revive usage in api.go and parse.go, indicating that package naming will not be changed for now.

* Update comments and dictionary entries for clarity

Added 'cunyue' to the project dictionary and updated inline comments in Go files to use proper comment syntax for clarity and consistency.

* Remove nolint directive from api package declaration

Deleted the 'nolint:revive' comment from the package declaration in api.go, allowing standard linting checks to apply.

* Remove unused api.go and update parse.go package comment

Deleted the core/internal/api/api.go file as it is no longer needed. Also updated the package comment in parse.go to remove the nolint directive.

* Add API package and update linter config

Introduces a new api.go file for the SwanLab API, updates the .golangci.yml to configure revive linter for 'API' variable naming, and adds a nolint directive in parse.go to suppress revive warnings for the package name.

* Remove unused comment and fix package declaration

Cleaned up the package declaration by removing the nolint directive and an extraneous 'revive' line in parse.go.

* Update revive linter rules in golangci config

Removes the custom var-naming rule for 'API' in the revive linter and adds a new exclusion for var-naming warnings in the internal/api directory.
* Refactor COS upload and experiment state APIs

Removed boto3/botocore dependencies and the CosClient, replacing direct COS upload logic with a new API-based approach using presigned URLs. Moved experiment state update logic to a dedicated API module and updated all usages. Refactored Client to remove COS logic, simplified session handling, and updated uploader and callback code to use the new APIs. Updated imports and tests to reflect file moves and API changes.

* Reset buffer pointer before COS upload

Added buffer.seek(0) before uploading to COS to ensure the buffer pointer is at the start. Also removed unused COS upload tests and related imports from test_client.py for code cleanup.

* Move buffer seek to _upload function

Relocated the buffer.seek(0) call from upload_to_cos to the _upload function to ensure the buffer is reset before uploading. This centralizes buffer handling and prevents redundant seeks.

* Add retry logic to file upload and unit tests

Enhanced the file upload function with retry logic and exponential backoff for robustness against network errors. Added comprehensive unit tests to verify upload success, retry behavior, exception handling, and server error scenarios.

* Improve error handling in COS upload function

Enhanced the upload_to_cos function to better handle exceptions during concurrent uploads. Failed uploads are now logged with warnings and retried, with errors on retry attempts also logged for improved reliability and debugging.

* Improve retry backoff in upload_file function

Changed the retry delay in upload_file to use exponential backoff (2^(attempt-1) seconds) instead of linear. Also removed a redundant comment in upload_to_cos for clarity.

* Remove unused import and URL prefix logic

Deleted the import of get_host_api and related code that handled URL prefixing for private environments, as it is no longer needed.
* Add empty input checks to uploader functions

Added early return statements to uploader functions in upload.py to handle empty input lists, preventing unnecessary processing and logging debug messages. Also updated comments and TODOs for consistency in batch.py and upload.py.

* Update upload.py

* Refactor upload functions with skip_if_empty decorator

Introduced a skip_if_empty decorator to reduce code duplication by handling empty input lists for upload functions. This change improves code readability and maintainability by centralizing the empty-check logic and debug logging.

* Refactor empty list checks in uploader functions

Simplified and unified empty list checks by removing redundant length checks in upload_files and upload_columns, and improving logic in skip_if_empty and upload_media_metrics. This streamlines the code and relies on existing mechanisms for handling empty inputs.

* Improve skip_if_empty decorator argument handling

Refactored the skip_if_empty decorator to use inspect for more robust detection of the first parameter, supporting both positional and keyword arguments. This ensures the decorator works correctly regardless of how the decorated function is called.
* Doc: Add the arxiv link of "MolAct" to the readme file

* remove the blank line

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
* feat: Add AMD ROCm GPU information collection functionality, supporting both Linux and Windows.

* fix: Improve AMD GPU information collection with error handling and fallback for amd-smi

* refactor: 移除 amd-smi 支持,统一使用 rocm-smi 进行 AMD GPU 信息采集

* fix: 增强 rocm-smi 的错误处理,确保在崩溃时不影响全局状态

* fix: 增强 rocm-smi 监控命令的错误处理,确保在崩溃时禁用后续监控

* refactor: 移除对 Windows 的支持,简化 AMD GPU 信息采集逻辑,仅保留对 Linux 的支持

* refactor: 优化 AMD GPU 信息采集逻辑,支持 Windows 和 Linux 系统,增强稳定性

* refactor: 优化 AMD GPU 信息采集逻辑,移除对 Windows 的支持,增强 Linux 系统的稳定性和性能

* chore: 更新文件时间戳为 2025-12-25

* docs: 添加对 AMD ROCm 的支持信息到 README 文件

* fix: 捕获异常时明确指定 Exception,增强代码健壮性

* update

* update README

---------

Co-authored-by: ZeYi Lin <[email protected]>
* support iluvatar gpu

* update link

* Update swanlab/data/run/metadata/hardware/gpu/iluvatar.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* update README

---------

Co-authored-by: Zhining Ma <[email protected]>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: ZeYi Lin <[email protected]>
* Add Timer utility with tests and update imports

Introduced a new Timer utility in swanlab.core_python.utils for scheduled and repeated task execution using daemon threads and event signaling. Added comprehensive unit tests for Timer, updated requirements-dev.txt to include pytest-mock, and adjusted __init__.py imports to expose the new utility.

* Refactor hardware monitor and remove unused helpers

Replaced MonitorCron with timer.Timer and a new monitor_interval function for hardware monitoring in SwanLabRun. Removed the check_log_level utility and its test, simplifying helper.py and cleaning up related test code. Also commented out problematic imports in __init__.py to avoid circular dependencies.

* Remove unused run lock from Timer class

Eliminated the threading.Lock previously used for mutual exclusion in the Timer class, as it was not necessary for task execution. The task execution method now directly calls the task without lock protection.
* Refactor client utilities and add experiment heartbeat

Moved client utility functions and models to a new utils.py file, including the safe_request decorator. Introduced a client heartbeat mechanism to keep experiments active, with integration in CloudPyCallback. Updated imports and improved type hinting for better maintainability.

* Add heartbeat interval parameter and improve cleanup

Added an interval parameter to create_client_heartbeat for configurable heartbeat timing. Improved CloudPyCallback.on_stop to check for heartbeat existence before canceling and joining, preventing potential errors.

* Fix typo in comment in client __init__.py

Corrected a typo in a comment to improve clarity and maintain code quality. No functional changes were made.
* fix

* fix: update return type annotation for map_amd_gpu_linux

Change return type from Tuple[Optional[str], dict] to Tuple[Optional[str], Optional[dict]] to match the (None, None) return case.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <[email protected]>

---------

Co-authored-by: Claude <[email protected]>
* able to modify session retry_times (disable the adapter when retries=0)

- Temporarily modify the retry count in the adapter by injecting the request count into the request headers.

* add unit_test for custom session retries

* add comment

* fix bug occurs in TimeoutHTTPAdapter when timeout=None

* handle ValueError caused by int(_retry)

- simplified some code in SwanSession

* Add retries parameter to client HTTP methods

Refactored Client to use SessionWithRetry and updated HTTP method signatures to accept an optional retries parameter. Overrode method signatures in SessionWithRetry to support retries and updated batch uploader to explicitly set retries to 0 for metric tracing. This improves control over request retry behavior.

* Validate retry count in TimeoutHTTPAdapter

Added a check to ensure the retry count is a non-negative integer in TimeoutHTTPAdapter. Updated error messages for clarity and added a unit test to verify ValueError is raised for invalid retry values.

---------

Co-authored-by: Kang Li <[email protected]>
…ter (SwanHubX#1421)

* feat: support tensorboard types filter and refactor converter

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <[email protected]>

* time sleep

* fix for gemini

---------

Co-authored-by: Claude <[email protected]>
* opt log user experience

- add timestamp for warning when network error occur
- Reduce the frequency of network error log printing

* accept gemini suggestions
* feat: treat empty run_id string as equivalent to None

* fix unit test

* fix unit test and make None check better readability
* Fix notification URL construction logic

* update

---------

Co-authored-by: ZeYi Lin <[email protected]>
* feat(init): api login (SwanHubX#1380)

* feat: get all projects through OpenApi (SwanHubX#1382)

* feat: get all project of a group

* update unit test using magicMock

* update unit test & fix bug

* opt import and rename some file

- move api from folder core_python to api

* fix wrong comment

* remove pending status when getting entity projects

* opt Project class

- add EN comment for Project properties
- add __str__ method for project label
- adapt snake_case naming for project properties
- change some project property name: projectLabels -> label
- change api.projects() param name: entity -> workspace
- delete some project properties (group, cuid)
- prase url inside Project class

* dynamically request project data according to the traversal of the project, rather than requesting all project information at once

* opt import

* iteratively request project data inside the Projects class

- move api form api package to core_python package

* fix bug & add type

* move api's type.py to core_python

* Feat/api runs user (SwanHubX#1403)

* feat: get runs metadata

- add api in core_python to get all exps in a project
- add new type and class to parse exps

* update unit test & fix bugs

* opt code style & add comment

* raise value error if user's path is invaded

* fix bug of Incorrectly raised ValueError

* get single exp through OpenApi

* resolve conflicts: filtering exp (basic implement)

- get full single exp info through filter func

# Conflicts:
#	swanlab/core_python/api/experiment.py

* opt __dict__ for project and experiment

* resolve conflict: update run.history()

* feat: update run.scan_history()

- add example codes

* resolve conflict: fix bugs when importing pandas

* import pandas only when used

* fix: pandas warning

* feat: run and user features of OpenApi

- get metric data in batch
- recover old OpenApi version for smooth transition
- create & delete api_key through Api.user

* opt Client importing in core_python's api

* opt logic of run.history()

- get latest api_keys when user delete api_key

* rename the deprecated openApi folder

* refactored run.history() & add return type

- add return type for the backend interfaces
- simplified get_experiment_metrics() and handle the csv inside HistoryPool

* removed redundant types

* update unit test using mock

* feat: get user's team

* feat: create user if user is the root user in self-hosted swanlab

- refactor the models in OpenApi

* feat: update unit test for create user and get user teams

* delete test code of OpenApi

* use ThreadPoolExecutor in HistoryPool

* Refactor Api class and cleanup api module structure

Moved the Api class implementation from swanlab/api/api.py to swanlab/api/__init__.py and deleted the redundant api.py file. Updated imports and references accordingly. Minor docstring and comment improvements, and fixed a message in thread.py to reference 'Api' instead of 'OpenApi'.

* Fix circular import by updating Api export

Combined the import of OpenApi and Api from .api to prevent circular import issues and simplify the export process.

---------

Co-authored-by: ZeYi Lin <[email protected]>
Co-authored-by: Kang Li <[email protected]>

* refactor open api (SwanHubX#1411)

* Refactor the OpenAPI codes, managing the code with a modular approach

* combine 3 kinds of user into one

- use @cached_property to cache some property of user

* Place the API's custom type into a separate module

- place class ApiBase into a single model

* get the user identity at the init of the OpenApi

- only the root user get by api.user() can create new account

* refactor module name & structure in api and core_python

* Modified parameter passing convention & add client property to ApiBase

- pass param through kwargs except client in OpenApi and core_python

* update code commit

- opt str() of Label class

* feat: open api unit tests (SwanHubX#1420)

* add unit test

- add unit test and fix bugs of getting all projects through OpenApi
- add unit tests for OpenApi runs and run.history()

* add unit test for api.user()

* accept suggestions from gemini

* fix bugs in projects

* revert changes

* Opt open api code (SwanHubX#1424)

* delete class ApiBase

* accept suggestions

- place yield inside the loop of the projects
- check if creating & deleting user or api_keys inside the api func
- add property is_self in user

* opt user api func

* fix bugs in projects

* accept gemini suggestions

- fix type error in self_hosted.py and user's api func
- add a constant for project page size

* opt over-encapsulated code

- discard constant PAGE_SIZE
- let adapter handle the exception when creating & deleting user info
- opt unit test for projects

* delete unused utils

* resolve conflict

* check self hosted info using a wrapper instead of checking when init (SwanHubX#1425)

* use a wrapper to check self_hosted info for some functions

- update unit tests for user (add self_hosted context)

* fix bug in OpenApi.list_workspaces()

* accept gemini suggestions

* raise value error when not being self_hosted

* raise value error when self_hosted is not available

- raise error when accessing unsupported user functions

* feat: workspace and json() (SwanHubX#1428)

* feat: replace the workspace field in the project object with a workspace object (SwanHubX#1430)

* Replace the workspace field in the project object with a workspace object

* cached the workspace object

* fix: fix bugs when getting info when 'profile' param is None (SwanHubX#1440)

* fix bugs when getting exp info when profile is None

- use getattr() to get profile and related info

* accept gemini suggestions

* feat: get all user when login as root user (SwanHubX#1438)

* get all user when login as root user

- opt @self_hosted decorator identity checking
- add api func get_users()

* accept gemini suggestions

* move get all users function from user to api

- opt identity checking logic

* Refactor user listing to use Users iterator class

Replaced the users() method in Api to return a new Users iterator class instead of yielding User objects directly. Added swanlab/api/users/__init__.py to encapsulate user iteration logic, improving code organization and separation of concerns.

* Add user pagination test and utility function

Added a test for paginated user retrieval in test_user.py and implemented create_user_data in utils.py to simulate paginated user data. Also clarified the docstring in create_user for self-hosted admin restriction.

---------

Co-authored-by: Kang Li <[email protected]>

* feat: get single project through openApi (SwanHubX#1439)

* get single project through openApi

- add api func get_project_info

* accept gemini suggestions

* Refactor workspace and project param names to 'path'

Replaces 'workspace' parameters with 'path' across API, core_python, and test modules for consistency. Updates related function calls, class initializations, and docstrings to reflect the new naming convention.

* Refactor Workspace initialization and usage

Updated Workspace class to require client and data parameters, removing path-based initialization logic. Refactored related API, Project, and Workspaces classes to fetch workspace data before instantiating Workspace, ensuring consistent and explicit data handling.

---------

Co-authored-by: Kang Li <[email protected]>

* Inline experiment history CSV fetch; remove pool

Remove the threaded HistoryPool helper and simplify Experiment.history by fetching per-key CSVs inline. Deleted swanlab/api/experiment/thread.py and removed its import. The history method now normalizes keys (allowing a single string), appends x_axis, deduplicates, retrieves CSV URLs via the client, reads them with pandas, strips a common prefix and trailing "_step" from column names, and concatenates DataFrames (inner join when x_axis is present, outer otherwise). When x_axis is given, timestamp columns are dropped, the x_axis column is validated and moved to the front. The method now raises if keys is None and supports sample trimming. Note: concurrency via HistoryPool is removed and the pandas parameter is effectively unused (the method returns a pandas DataFrame).

* Rename history() to metrics() in Experiment

Remove Experiment.config and Experiment.summary properties and rename the Experiment.history(...) method to Experiment.metrics(...). Update docstrings, example usage and the pandas import error message to reference metrics. This is an API change — update callers to use exp.metrics(...) and note that config/summary accessors were removed.

* Refactor DataFrame joining and x_axis handling

Introduce x_axis_state to centralize x_axis checks and only append x_axis to keys when applicable. Replace pd.concat(join_type) with iterative outer .join(...).sort_index() to build the result DataFrame, and keep timestamp columns dropped as before. After reordering columns to put x_axis first, filter out rows where x_axis is NaN. Minor cleanup and readability improvements.

* fix test

* fix test

* Improve Experiment.metrics CSV parsing, validation

Refactor Experiment.metrics: tighten keys validation (keys must be a non-empty list of strings), simplify docstring and pandas import error message, and treat pandas param as reserved. Normalize x_axis handling (default 'step'), append x_axis when needed, and fetch each metric CSV. Extract common prefix from the first column, strip that prefix and the "_step" suffix in a Python 3.8-compatible way, then outer-join and sort the DataFrames. When x_axis is used, drop timestamp columns, ensure x_axis exists, move it to the first column and drop rows with null x_axis. Finally apply the optional sample limit.

---------

Co-authored-by: Bainianzzz <[email protected]>
Co-authored-by: ZeYi Lin <[email protected]>
Co-authored-by: Bainianzzz <[email protected]>
Replace ClickHouse-based metric storage with a fixed-length binary format (.swd)
backed by object storage (COS AppendObject). This enables O(1) random access,
HTTP Range GET incremental reads, and crash recovery via tail truncation.

Phase 1 - Binary format: 128-byte header + 24-byte scalar / 80-byte media index
          records with SwdWriter, SwdReader, and format utilities.
Phase 2 - Local storage: ManifestManager, feature flags (use_swd_format,
          use_cos_direct), dual-write JSONL+.swd in SwanLabKey and callbacks.
Phase 3 - COS integration: CosClient, CosAppender with position verification,
          exponential backoff retry, CloudPyCallback COS lifecycle.
Phase 4 - Manifest lifecycle: periodic COS sync timer, ConsoleLogAppender
          with 4KB buffered AppendObject upload.
Phase 5 - Backward compatibility: BackupToSwdConverter, JsonlToSwdConverter,
          SwdExporter (CSV/JSONL/Parquet), sync command .swd detection.

Includes 98 unit tests and swd-format.md specification document.

Co-Authored-By: Claude Opus 4.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.