CephFS与RGW实战及性能压测

CephFS与RGW实战及性能压测
[TOC]
为什么需要CephFS
RBD 只能供单个客户端挂载——同一块镜像不能被两台主机同时读写,CephFS 解决的就是 “多节点共享同一份数据” 的问题
| RBD | CephFS | |
|---|---|---|
| 共享能力 | 单客户端独占 | ✅ 多客户端同时读写 |
| 依赖组件 | 不需要 MDS | 需要 MDS |
| 适用场景 | 虚拟机磁盘、数据库 | 文件共享、日志 |
| 存储类型 | 接口(接口层) | 工具(工具层) | 依赖关系 | 通俗理解 | 典型用途 |
|---|---|---|---|---|---|
| 块存储 | ==RBD== | rbd | 依赖 librados | 虚拟硬盘,能格式化、分区、挂载 | 虚拟机磁盘、数据库 |
| 文件存储 | ==CephFS== | ceph fs | 不依赖 librados | 共享文件夹,多台机器同时挂载 | 文件共享、日志存储 |
| 对象存储 | ==RadosGW== | radosgw-admin | 依赖 librados | HTTP/HTTPS API,适合非结构化数据 | 图片/视频云存储 |
RBD ──────> librados ──> RADOS ClusterRadosGW ──> librados ──> RADOS ClusterCephFS ────────────────> RADOS Cluster (直接和 MDS / OSD 通信)- CephFS 底层走 RADOS,需要一个额外的守护进程——MDS(Metadata Server)
- 它只维护内存索引,不存任何实际数据

一个文件读写的完整路径:
客户端发起读写→ 1. 问 MDS:对应的 inode 索引→ 2. 文件切成 4MB 对象 → CRUSH 哈希 → 分散命中多个 PG→ 3. 每个 PG 按 size=3 副本落到不同 OSD('多磁盘并行读写')→ 4. OSD 返回数据 → 客户端拼回完整文件为什么比 NFS 快:NFS 所有读写挤在服务端单块盘上,CephFS 同时读写多块 OSD
- CephFS由多个存储节点组成,扩展性相当的高
- CephFS 底层 RADOS 自带数据冗余,无单点故障
- NFS 服务端本身是单点,需要额外的高可用方案
部署CephFS
创建存储池并部署MDS
root@Ceph01 ~# ceph osd pool create opsfs_data--autoscale_mode # PG自动伸缩默认是开启的~# 一个是数据池,另一个是元数据池root@Ceph01 ~# ceph osd pool create opsfs_metadataroot@Ceph01 ~# ceph osd pool set opsfs_data bulk true # 标记为大容量池(PG自动伸缩)root@Ceph01 ~# ceph osd pool ls detail | grep opsfs`需要等一会才会出结果!`pool 7 'opsfs_data' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 256 pgp_num 129 pgp_num_target 256 autoscale_mode on last_change 264 lfor 0/0/264 flags hashpspool,bulk stripe_width 0 read_balance_score 2.04# pg_num 256 pgp_num 129 (延迟30s)pool 8 'opsfs_metadata' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 262 lfor 0/0/260 flags hashpspool stripe_width 0 read_balance_score 2.81
# 创建一个文件系统ceph fs new <文件系统名> <元数据池> <数据池> ↓ ↓ ↓ ↓ 新建 起个名 存文件索引 存文件内容root@Ceph01 ~# ceph fs new ops-cephfs opsfs_metadata opsfs_data
`ceph orch apply mds <fs-name>`# "让编排器给 ops-cephfs 这个文件系统拉起 MDS 进程"root@Ceph01 ~# ceph orch apply mds ops-cephfs 1. 会启动两个 MDS 守护进程(分别在两台主机上) 2. 一个设为 active,一个设为 standby;与 MGR 的高可用结构一致 3. MDS 连到 opsfs_metadata 池,开始维护 inode 索引验证MDS状态
# 看 MDS 进程整体概况(所有)root@Ceph01 ~# ceph mds statops-cephfs:1 {0=ops-cephfs.Ceph02.xlaleh=up:active} 1 up:standby
# 查看某个文件系统的详细信息root@Ceph01 ~# ceph fs status ops-cephfsops-cephfs - 0 clients==========RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 active ops-cephfs.Ceph02.xlaleh Reqs: 0 /s 10 13 12 0 POOL TYPE USED AVAILopsfs_metadata metadata 157k 56.5G opsfs_data data 0 56.5G STANDBY MDSops-cephfs.Ceph01.rltckvMDS version: ceph version 20.2.2 (0f...c) tentacle (stable)# Ceph02活跃,Ceph01备用
root@Ceph01 ~# ceph fs lsname: ops-cephfs, metadata pool: opsfs_metadata, data pools: [opsfs_data ]
root@Ceph01 ~# ceph -s | grep mds mds: 1/1 daemons up, 1 standbyMDS高可用验证
root@Ceph01 ~# ceph fs status ops-cephfs | egrep -A1 "active|STANDBY" 0 active ops-cephfs.Ceph02.xlaleh # active 在 Ceph02 STANDBY MDS ops-cephfs.Ceph01.rltckv # standby 在 Ceph01
# 关掉 Ceph02 模拟故障root@Ceph02 ~# init 0
# 等待约 30s,standby 自动顶上root@Ceph01 ~# ceph fs status ops-cephfs | grep active 0 active ops-cephfs.Ceph01.rltckv # replay → active(秒级切换)客户端挂载(内核方式)
Ubuntu 24.04 内核自带 ceph 模块,直接 mount -t ceph 即可,无需安装额外包
创建用户并导出key
# Ceph01 上操作root@Ceph01 ~# ceph auth add client.cephfs mon 'allow r' mds 'allow rw' osd 'allow rwx'root@Ceph01 ~# ceph auth add client.cephfs-ro mon 'allow r' mds 'allow r' osd 'allow r'# 只读用户root@Ceph01 ~# ceph auth print-key client.cephfs ;echoAQA5wGRqh4c+IRAAw4a9Nh3Ukek0A4ziYVgREg==root@Ceph01 ~# ceph auth get-or-create-key client.cephfs-roAQBBwGRqWp0HBRAALDOXr+Zvj7TnVoQCOUBWwA=='两种获取用户key的方式'挂载并写入
# Ceph02 上操作root@Ceph02 ~# mkdir -p /mnt/cephfsroot@Ceph02 ~# mount -t ceph 10.0.0.3:6789,10.0.0.4:6789,10.0.0.5:6789:/ /mnt/cephfs -o name=cephfs,secret=AQA5wGRqh4c+IRAAw4a9Nh3Ukek0A4ziYVgREg==# -o name=用户名,secret=key值......(2) No such file or directory# 虽有报错,实则已经挂在上去了~root@Ceph02 ~# df -h | grep cephfs10.0.0.3:6789,10.0.0.4:6789,10.0.0.5:6789:/ 57G 0 57G 0% /mnt/cephfsroot@Ceph02 ~# cp /etc/hosts /etc/fstab /mnt/cephfs/root@Ceph02 ~# ls /mnt/cephfs/fstab hosts# 该cephfs用户可读写多客户端共享
# Ceph03 同样挂载root@Ceph03 ~# mkdir -p /mnt/cephfsroot@Ceph03 ~# mount -t ceph 10.0.0.3:6789,10.0.0.4:6789,10.0.0.5:6789:/ /mnt/cephfs -o name=cephfs,secret=AQA5wGRqh4c+IRAAw4a9Nh3Ukek0A4ziYVgREg==# 这个也是可读写的用户root@Ceph03 ~# df -h | grep cephfs10.0.0.3:6789,10.0.0.4:6789,10.0.0.5:6789:/ 57G 0 57G 0% /mnt/cephfs
# Ceph03 写入root@Ceph03 ~# cp /etc/os-release /mnt/cephfs/root@Ceph03 ~# ls /mnt/cephfs/fstab hosts os-release
# Ceph02 & Ceph03 数据一致(内容共享)root@Ceph02 ~# ls /mnt/cephfs/fstab hosts os-release权限隔离
root@Ceph03 ~# mkdir /mnt/cephfs-ro
# Ceph03 上用只读用户挂载root@Ceph03 ~# mount -t ceph 10.0.0.3:6789,10.0.0.4:6789,10.0.0.5:6789:/ /mnt/cephfs-ro -o name=cephfs-ro,secret=AQBBwGRqWp0HBRAALDOXr+Zvj7TnVoQCOUBWwA==
root@Ceph03 ~# df -h | grep cephfs10.0.0.3:6789,10.0.0.4:6789,10.0.0.5:6789:/ 57G 0 57G 0% /mnt/cephfs10.0.0.3:6789,10.0.0.4:6789,10.0.0.5:6789:/ 57G 0 57G 0% /mnt/cephfs-roroot@Ceph03 ~# ls /mnt/cephfsfstab hosts os-release # 读写root@Ceph03 ~# ls /mnt/cephfs-rofstab hosts os-release #⚠️ 只读root@Ceph03 ~# cp /etc/group /mnt/cephfs-ro/ # ❌ Permission deniedcp: cannot create regular file '/mnt/cephfs-ro/group': Permission deniedroot@Ceph03 ~# cp /etc/group /mnt/cephfs # ✅ 读写挂载cephfs指定的路径
# 以Ceph02为例root@Ceph02 ~# tree /mnt/cephfs//mnt/cephfs/├── fstab├── group├── hosts└── os-releaseroot@Ceph02 ~# mkdir -p /mnt/cephfs/dataroot@Ceph02 ~# mkdir -p /mnt/cephfs/resourcesroot@Ceph02 ~# touch /mnt/cephfs/data/index{1..2}.htmlroot@Ceph02 ~# touch /mnt/cephfs/resources/test{1..2}.phproot@Ceph02 ~# tree /mnt/cephfs//mnt/cephfs/├── data│ ├── index1.html│ └── index2.html├── fstab├── group├── hosts├── os-release└── resources ├── test1.php └── test2.php
# 我们只去挂载根下的/dataroot@Ceph02 ~# mkdir /mnt/pathroot@Ceph02 ~# mount -t ceph 10.0.0.3:6789,10.0.0.4:6789,10.0.0.5:6789:/data /mnt/path -o name=cephfs,secret=AQA5wGRqh4c+IRAAw4a9Nh3Ukek0A4ziYVgREg==root@Ceph02 ~# ls /mnt/path/index1.html index2.html# data目录下只有html页面自动挂载
CephFS 不需要 rbd map,直接 mount -t ceph 即可——用 systemd 管理开机自启。
1)编写挂载脚本⚠️ 分隔符必须写成 <<'EOF',否则 ${} 会被 bash 提前展开root@Ceph02 ~# cat > /usr/local/bin/cephfs-auto-mount.sh <<'EOF'#!/bin/bash
MON_ADDR="10.0.0.3:6789,10.0.0.4:6789,10.0.0.5:6789"KEY="AQA5wGRqh4c+IRAAw4a9Nh3Ukek0A4ziYVgREg=="MNT="/mnt/cephfs"
if mountpoint -q "${MNT}"; then echo "${MNT} 已挂载,跳过" exit 0fi
mkdir -p "${MNT}"mount -t ceph "${MON_ADDR}:/" "${MNT}" -o name=cephfs,secret=${KEY}echo "CephFS → ${MNT} 挂载完成"EOFroot@Ceph02 ~# chmod +x /usr/local/bin/cephfs-auto-mount.sh
2)编写 systemd 服务单元root@Ceph02 ~# cat > /etc/systemd/system/cephfs-mount.service <<EOF[Unit]Description=CephFS Auto MountAfter=network-online.target# 确保网络就绪后再执行(Ceph 集群通信依赖网络)
[Service]Type=oneshot# 一次性任务,执行完即结束ExecStart=/usr/local/bin/cephfs-auto-mount.shRemainAfterExit=yes# 执行完后仍视为 active,方便追踪状态
[Install]WantedBy=multi-user.targetEOF
3)启用并验证root@Ceph02 ~# systemctl daemon-reloadroot@Ceph02 ~# systemctl enable cephfs-mount.serviceCreated symlink /etc/systemd/system/multi-user.target.wants/cephfs-mount.service ...root@Ceph02 ~# systemctl start cephfs-mount.serviceroot@Ceph02 ~# df -h | grep cephfs10.0.0.3:6789,...:/ 57G 0 57G 0% /mnt/cephfs
4)重启验证root@Ceph02 ~# rebootroot@Ceph02 ~# df -h | grep cephfs10.0.0.3:6789,...:/ 57G 0 57G 0% /mnt/cephfs✅ 开机自动挂载成功
5)查看日志root@Ceph02 ~# systemctl status cephfs-mount.service● cephfs-mount.service - CephFS Auto Mount Active: active (exited)root@Ceph02 ~# journalctl -u cephfs-mount | tail -2CephFS → /mnt/cephfs 挂载完成基于用户空间FUSE方式访问
内核低于 4.x 或没有 ceph 模块时,可用 ceph-fuse 替代——工作在用户空间,性能不如内核挂载
- ⚠️ 老旧内核才需要
ceph-fuse
# 查看linux内核是否有这个ceph模块root@Ceph02 ~# modinfo cephfilename: /lib/modules/6.8.0-136-generic/kernel/fs/ceph/ceph.ko.zstlicense: GPLdescription: Ceph filesystem for Linuxauthor: Patience Warnick <patience@newdream.net>author: Yehuda Sadeh <yehuda@hq.newdream.net>.................# 查看正在启用的内核模块(已经挂载上CephFS)root@Ceph02 ~# lsmod | grep cephceph 626688 1libceph 548864 1 cephnetfs 512000 1 cephlibcrc32c 12288 6 nf_conntrack,nf_nat,btrfs,nf_tables,raid456,libceph
# 1. 安装 fuse 客户端工具apt -y install ceph-fuse
# 2. 服务端导出用户 keyringceph auth export client.cephfs -o ceph.client.cephfs.keyringscp ceph.client.cephfs.keyring root@10.0.0.4:/tmp/
# 3. 客户端挂载mkdir -p /mnt/cephfsceph-fuse -n client.cephfs -m 10.0.0.3:6789,10.0.0.4:6789,10.0.0.5:6789/ /mnt/cephfs \ -c /tmp/ceph.client.cephfs.keyring# -n 用户名 -m Monitor地址 -c 指定keyring文件位置删除CephFS
# 删除 MDS 编排规则(停 MDS 容器)root@Ceph01 ~# ceph orch rm mds.ops-cephfs --force
# 标记文件系统为不可用(阻止 MDS 再加入)root@Ceph01 ~# ceph fs fail ops-cephfs
# 删除文件系统定义root@Ceph01 ~# ceph fs rm ops-cephfs --yes-i-really-mean-itroot@Ceph01 ~# ceph fs lsNo filesystems enabledroot@Ceph01 ~# ceph osd pool ls | grep opsfsopsfs_dataopsfs_metadata# 这些池子还没有删除对象存储网关 RGW
网关概述
📌 一句话:RGW(RadosGW) 是 Ceph 的对象存储网关,把 Ceph 集群包装成一套兼容 S3 / Swift 的 HTTP 接口,让应用用 PUT/GET 就能存对象,而不用关心底层的 RADOS
Ceph Client ──HTTP(S)──> RadosGW(Civetweb) ──librados──> RADOS Cluster (S3/Swift API) (rgw 守护进程) (OSD 存真实数据)- RGW 守护进程叫
rgw,底层走librados,和 RBD 一样直接读写 RADOS - 默认内置 Civetweb 作为 Web Server(Squid 版起默认用 80 端口,老版本用 7480)
- 提供两类兼容接口:
- Amazon S3:
User(用户) -> Bucket(桶) -> Object(对象)三层结构- 桶从属于用户,用户名即桶的命名空间(不同用户的桶名可重复)
- OpenStack Swift:
Account(账户) -> container(容器) -> Object(对象)- User 拥有整个账户,Subuser 仅拥有指定容器的访问权限
- Amazon S3:
- RGW 自己有一套用户体系(
radosgw-admin user create),和集群的ceph auth是两套
`OpenStack Swift`Account(账户) ├── User(主用户) └── Subuser(子用户)# User 和 Subuser 是身份凭证,不参与 URL 路径# Subuser 的权限可以精细到某个具体的 Container💡 RGW 首次启动时会自动创建一批以 default. 为前缀的存储池(如 default.rgw.meta、default.rgw.buckets.index、default.rgw.buckets.data 等),不用手工建池
但在生产环境中,建议手动预先创建并调优这些池,以确保性能和稳定性
部署 RGW
1)查看集群现状root@Ceph01 ~# ceph -s cluster: id: a9b741cd-856a-11f1-b216-525400c6b9fa health: HEALTH_OK services: mon: 3 daemons, quorum Ceph01,Ceph02,Ceph03 (age 70s) [leader: Ceph01] mgr: Ceph01.stlhbi(active, since 59s), standbys: Ceph02.lbqijv osd: 9 osds: 9 up (since 59s), 9 in (since 4w)'确认集群健康,9 个 OSD 已就位'
2)部署 rgw 服务(服务名用 kpyun,落 Ceph01/Ceph02 两台,端口 80)root@Ceph01 ~# ceph orch apply rgw kpyun --placement="2 Ceph01 Ceph02" --port 80Scheduled rgw.kpyun update...'cephadm 会自动拉起容器、建池、注册服务'
3)查看 rgw 组件是否起来root@Ceph01 ~# ceph -s services: mon: 3 daemons, quorum Ceph01,Ceph02,Ceph03 (age 70s) [leader: Ceph01] mgr: Ceph01.stlhbi(active, since 59s), standbys: Ceph02.lbqijv osd: 9 osds: 9 up (since 59s), 9 in (since 4w) rgw: 2 daemons active (2 hosts, 1 zones) # 👈 出现 rgw 了
4)看 rgw 守护进程详情root@Ceph01 ~# ceph orch ps --service_name rgw.kpyunNAME HOST PORTS STATUS REFRESHED AGE MEM USE VERSIONrgw.kpyun.Ceph01.saocsj Ceph01 *:80 running (33s) 23s ago 33s 121M 20.2.2rgw.kpyun.Ceph02.zjvuxs Ceph02 *:80 running (30s) 24s ago 30s 120M 20.2.2'两台各一个 rgw 容器,都监听 80'
5)看 rgw 自动建的池root@Ceph01 ~# ceph osd pool ls | grep -i rgw.rgw.rootdefault.rgw.logdefault.rgw.controldefault.rgw.meta'就这 4 个,object 数据池(default.rgw.buckets.data)等用到时才出现'
6)访问 WebUI(HTTP 接口)root@Ceph01 ~# curl -s -I http://10.0.0.3/ | head -1HTTP/1.1 200 OK✅️ 返回 200
http://10.0.0.3/http://10.0.0.4/
s3cmd 操作对象存储
初始化配置
1)安装 s3cmd 工具包(Ubuntu 用 apt)root@Ceph01 ~# apt -y install s3cmdroot@Ceph01 ~# s3cmd --versions3cmd version 2.4.0'装好,版本 2.4.0'
2)创建 rgw 账号(在 Ceph01 上用 radosgw-admin)root@Ceph01 ~# radosgw-admin user create --uid "jiuzhao" --display-name "jzops"--uid # 用户ID(系统唯一标识) ❌ 全局唯一,不可重复--display-name # 显示名称(昵称/别名) ✅ 可以重复,建议唯一{ "user_id": "jiuzhao", "display_name": "jzops", "email": "", "suspended": 0, "max_buckets": 1000, "subusers": [], "keys": [ { "user": "jiuzhao", "access_key": "42GA8G0I7JGTVGX43NPK", "secret_key": "8pJrGe843L6C8xemXAi3RhA2OE28WIguFcjKg247", "active": true, "create_date": "2026-08-25T10:29:15.951471Z" } ], ...}`拿到 access_key 和 secret_key,后面 s3cmd 要用`
3)写 s3cmd 配置文件 /root/.s3cfgroot@Ceph01 ~# cat > /root/.s3cfg <<'EOF'[default]access_key = 42GA8G0I7JGTVGX43NPKsecret_key = 8pJrGe843L6C8xemXAi3RhA2OE28WIguFcjKg247host_base = 10.0.0.3host_bucket = 10.0.0.3/%(bucket)use_https = FalseEOF'host_base 直接填 rgw 的 IP(10.0.0.3),不配域名也能用''host_bucket 用 路径风格: 10.0.0.3/%(bucket),因为没配 rgw dns name'
4)测试连通性root@Ceph01 ~# s3cmd ls'现在还没有 bucket,所以没输出,但只要不报错就说明 key 对了'✅️ 配置成功💡 也可以用 s3cmd --configure 交互生成 .s3cfg,本质就是上面这份配置,直接手写更可控
root@Ceph01 ~# s3cmd --configure
Enter new values or accept defaults in brackets with Enter.Refer to user manual for detailed description of all options.
Access key and Secret key are your identifiers for Amazon S3. Leave them empty for using the env variables.Access Key [42GA8G0I7JGTVGX43NPK]: # rgw账号的access_keySecret Key [8pJrGe843L6C8xemXAi3RhA2OE28WIguFcjKg247]: # rgw账号的secret_keyDefault Region [US]: # 直接回车即可
Use "s3.amazonaws.com" for S3 Endpoint and not modify it to the target Amazon S3.S3 Endpoint [10.0.0.3]: # 用于访问rgw的地址
Use "%(bucket)s.s3.amazonaws.com" to the target Amazon S3. "%(bucket)s" and "%(location)s" vars can be usedif the target S3 system supports dns based buckets.DNS-style bucket+hostname:port template for accessing a bucket [10.0.0.3/%(bucket)]: # 设置DNS解析风格
Encryption password is used to protect your files from readingby unauthorized persons while in transfer to S3Encryption password: # 文件不加密,直接回车即可Path to GPG program [/usr/bin/gpg]: # 指定自定义的gpg程序路径,直接回车即可
When using secure HTTPS protocol all communication with Amazon S3servers is protected from 3rd party eavesdropping. This method isslower than plain HTTP, and can only be proxied with Python 2.7 or newerUse HTTPS protocol [No]: # 你的rgw是否是https,如果不是设置为No
On some networks all internet access must go through a HTTP proxy.Try setting it here if you can't connect to S3 directlyHTTP Proxy server name: # 代理服务器的地址,我并没有配置代理服务器,因此直接回车即可
New settings: # 注意,下面的信息是上面咱们填写时一个总的预览信息 Access Key: 42GA8G0I7JGTVGX43NPK Secret Key: 8pJrGe843L6C8xemXAi3RhA2OE28WIguFcjKg247 Default Region: US S3 Endpoint: 10.0.0.3 DNS-style bucket+hostname:port template for accessing a bucket: 10.0.0.3/%(bucket) Encryption password: Path to GPG program: /usr/bin/gpg Use HTTPS protocol: False HTTP Proxy server name: HTTP Proxy server port: 0
Test access with supplied credentials? [Y/n] Y # 如果确认上述信息没问题的话,则输入字母Y即可Please wait, attempting to list all buckets...Success. Your access key and secret key worked fine :-)
Now verifying that encryption works...Not configured. Never mind.
Save settings? [y/N] y # 保存配置Configuration saved to '/root/.s3cfg'基础操作实战
1)创建一个新的存储桶 bucketroot@Ceph01 ~# s3cmd mb s3://kpyun-bucket# mb: 这是 s3cmd 的一个子命令,是 "make bucket" 的缩写,专门用来创建新的存储桶Bucket 's3://kpyun-bucket/' createdroot@Ceph01 ~# s3cmd ls2026-08-29 02:21 s3://kpyun-bucket✅️ bucket 建好了
2)准备测试文件并上传root@Ceph01 ~# ls /tmp | egrep 'bili|png'bilibili.webm `视频`s.png `图片`root@Ceph01 ~# echo "hello kpyun ceph storage" > /tmp/s3test.txtroot@Ceph01 ~# s3cmd put /tmp/s3test.txt /tmp/{bilibili.webm,s.png} s3://kpyun-bucketpload: '/tmp/s3test.txt' -> 's3://kpyun-bucket/s3test.txt'upload: '/tmp/bilibili.webm' -> 's3://kpyun-bucket/bilibili.webm'upload: '/tmp/s.png' -> 's3://kpyun-bucket/s.png'===================================`s3cmd 使用教程`root@Ceph01 ~# s3cmd --helpCommands: Make bucket s3cmd mb s3://BUCKET Remove bucket ✅ '删除整个存储桶及其所有元数据' s3cmd rb s3://BUCKET # 先递归删除桶里所有文件,再删除桶本身:s3cmd rb -r s3://kpyun-backup `桶被彻底删除,无法恢复` List objects or buckets ✅ '列出所有存储桶(不带路径时)' s3cmd ls [s3://BUCKET[/PREFIX]] List all object in all buckets ✅ '列出所有存储桶中的所有对象(文件)' s3cmd la # s3cmd la 命令的后面不应该再跟具体的存储桶路径;正确的用法应该是直接执行 s3cmd la # 要准确检查 某个存储通 是否已被清空,你应该使用专门针对单个存储桶的 ls 命令,而不是 la Put file into bucket ✅ '也可以是多个文件' s3cmd put FILE [FILE...] s3://BUCKET[/PREFIX] Get file from bucket ✅ '从 S3 存储桶下载文件到本地' s3cmd get s3://BUCKET/OBJECT LOCAL_FILE Disk usage by buckets ✅ '查看 存储桶/对象 磁盘使用量' s3cmd du [s3://BUCKET[/PREFIX]] Delete file from bucket ✅ `从 S3 存储桶中删除指定文件` s3cmd del s3://BUCKET/OBJECT Delete file from bucket (alias for del) ✅ `从 S3 存储桶中删除指定文件` s3cmd rm s3://BUCKET/OBJECT '两者等价,rm 是 del 的别名(alias),执行效果一模一样' # 删除所有文件:s3cmd rm s3://kpyun-backup/* `桶变为空桶,但桶本身还在,可以继续往这个桶里上传新文件`
3)下载回来验证内容一致root@Ceph01 ~# s3cmd get s3://kpyun-bucket/s3test.txt ./download: 's3://kpyun-bucket/s3test.txt' -> './s3test.txt'root@Ceph01 ~# cat ./s3test.txthello kpyun ceph storage'内容一致!'
4)查看 bucket 内容 / 对象信息 / 大小root@Ceph01 ~# s3cmd la2026-08-29 02:37 2053779 s3://kpyun-bucket/bilibili.webm2026-08-29 02:37 582512 s3://kpyun-bucket/s.png2026-08-29 02:34 25 s3://kpyun-bucket/s3test.txt
root@Ceph01 ~# s3cmd info s3://kpyun-bucket/s3test.txts3://kpyun-bucket/s3test.txt (object): File size: 25 Last mod: Sat, 29 Aug 2026 02:34:17 GMT MIME type: text/plain Storage: STANDARD MD5 sum: 8ddbf928658ad38b5ffdcf5618ae8d28 SSE: none Policy: none CORS: none ACL: jzops: FULL_CONTROL
root@Ceph01 ~# s3cmd du --human-readable s3://kpyun-bucket--human-readable = -H # 以人类可读格式显示 2M 3 objects s3://kpyun-bucket/`这里是查看整个存储桶大小,可以查看 单个文件/某个路径下所有文件 的大小`5)大文件分片上传(默认 15MB 一片)root@Ceph01 ~# dd if=/dev/zero of=/tmp/big.bin bs=1M count=40root@Ceph01 ~# ls -lh /tmp/big.bin-rw-r--r-- 1 root root 40M Aug 29 11:05 /tmp/big.binroot@Ceph01 ~# s3cmd put /tmp/big.bin s3://kpyun-bucketupload: '/tmp/big.bin' -> 's3://kpyun-bucket/big.bin' [part 1 of 3, 15MB🔥] [1 of 1]upload: '/tmp/big.bin' -> 's3://kpyun-bucket/big.bin' [part 2 of 3, 15MB🔥] [1 of 1]upload: '/tmp/big.bin' -> 's3://kpyun-bucket/big.bin' [part 3 of 3, 10MB🔥] [1 of 1]'40M 文件被自动按 15MB 切片上传'
6)桶间拷贝 / 移动 / 目录同步root@Ceph01 ~# s3cmd mb s3://kpyun-backupBucket 's3://kpyun-backup/' createdroot@Ceph01 ~# s3cmd ls | grep kpyun-backup2026-08-29 03:08 s3://kpyun-backup
root@Ceph01 ~# s3cmd cp s3://kpyun-bucket/big.bin s3://kpyun-backupremote copy: 's3://kpyun-bucket/big.bin' -> 's3://kpyun-backup/big.bin' [1 of 1]
root@Ceph01 ~# s3cmd mv s3://kpyun-bucket/s3test.txt s3://kpyun-backupmove: 's3://kpyun-bucket/s3test.txt' -> 's3://kpyun-backup/s3test.txt' [1 of 1]root@Ceph01 ~# s3cmd ls -H s3://kpyun-backup-H = --human-readable # 用短选项!2026-08-29 03:08 40M s3://kpyun-backup/big.bin2026-08-29 03:13 25 s3://kpyun-backup/s3test.txt
root@Ceph01 ~# s3cmd sync /var/log/ --exclude '*' --include 'auth.*' --dry-run s3://kpyun-backup# 默认情况下,所有文件都会被包含(相当于隐含了一个 --include '*')-n = --dry-run # 模拟运行(只显示会做什么,不实际执行)✅ 正确的做法:先用 --exclude 排除所有,再用 --include 挑选, --dry-run 干跑一遍exclude: ......... `排除了很多文件`upload: '/var/log/auth.log' -> 's3://kpyun-backup/auth.log'upload: '/var/log/auth.log.1' -> 's3://kpyun-backup/auth.log.1'upload: '/var/log/auth.log.2.gz' -> 's3://kpyun-backup/auth.log.2.gz'WARNING: Exiting now because of --dry-runroot@Ceph01 ~# s3cmd sync /var/log/ --exclude '*' --include 'auth.*' s3://kpyun-backupupload: '/var/log/auth.log' -> 's3://kpyun-backup/auth.log' [1 of 3]upload: '/var/log/auth.log.1' -> 's3://kpyun-backup/auth.log.1' [2 of 3]upload: '/var/log/auth.log.2.gz' -> 's3://kpyun-backup/auth.log.2.gz' [3 of 3]
root@Ceph01 ~# s3cmd ls -H s3://kpyun-bucket2026-08-29 03:05 40M s3://kpyun-bucket/big.bin2026-08-29 02:37 2005K s3://kpyun-bucket/bilibili.webm2026-08-29 02:37 568K s3://kpyun-bucket/s.pngroot@Ceph01 ~# s3cmd ls -H s3://kpyun-backup2026-08-29 03:45 64K s3://kpyun-backup/auth.log2026-08-29 03:45 3K s3://kpyun-backup/auth.log.12026-08-29 03:45 26K s3://kpyun-backup/auth.log.2.gz2026-08-29 03:38 40M s3://kpyun-backup/big.bin2026-08-29 03:38 25 s3://kpyun-backup/s3test.txt'cp/mv/sync 都正常,跨桶复制秒级完成'给 bucket 加匿名读策略,外部不用 key 也能下载对象(常用于 CDN / 静态网站)
1)写策略文件(允许任何人 GetObject)root@Ceph01 ~# cat > /root/anon-policy.json <<'EOF'{ // 版本号:策略语言的版本,固定用 "2012-10-17",这是AWS IAM策略的最终版本,不用改 "Version": "2012-10-17",
// 策略语句数组:可以包含多条规则 "Statement": [ { // Effect: 效果,Allow 表示"允许",Deny 表示"拒绝" "Effect": "Allow",
// Principal: 授权主体,{"AWS": ["*"]} 表示"允许所有人(匿名用户)" // 这里的 "*" 是通配符,代表所有AWS用户(包括未认证的) "Principal": {"AWS": ["*"]},
// Action: 允许执行的操作,"s3:GetObject" 表示"下载/读取对象" // 其他常见操作:s3:PutObject(上传)、s3:DeleteObject(删除)、s3:ListBucket(列出) "Action": "s3:GetObject",
// Resource: 资源范围,指定这个策略应用到哪些对象上 // "arn:aws:s3:::kpyun-bucket/*" 表示 kpyun-bucket 这个桶里的所有对象 // 注意:arn:aws:s3::: 是AWS的资源命名规范,固定写法 "Resource": ["arn:aws:s3:::kpyun-bucket/*"] } ]}EOF相关重要权限(写在 Action 里):
官方文档:https://docs.ceph.com/en/tentacle/radosgw/bucketpolicy/
| 权限 | 作用 |
|---|---|
s3:GetObject | 下载 / 读取对象 |
s3:PutObject | 上传 / 覆盖对象 |
s3:DeleteObject | 删除对象 |
s3:ListBucket | 列出 bucket 内容 |
s3:CreateBucket | 创建 bucket |
s3:DeleteBucket | 删除 bucket |
2)应用策略root@Ceph01 ~# s3cmd info s3://kpyun-bucket/ | grep -i policy Policy: noneroot@Ceph01 ~# s3cmd setpolicy /root/anon-policy.json s3://kpyun-bucket# 把这个策略文件应用(设置)到 kpyun-bucket 这个存储桶上s3://kpyun-bucket/: Policy updatedroot@Ceph01 ~# s3cmd info s3://kpyun-bucket/ | grep -A3 -i policy Policy: { // 版本号:策略语言的版本,固定用 "2012-10-17",这是AWS IAM策略的最终版本,不用改 "Version": "2012-10-17", ...... # 现在就有自己的策略了
3)匿名访问验证(这次不带 key)http://10.0.0.3/kpyun-bucket/bilibili.webm # 视频http://10.0.0.3/kpyun-bucket/s.png # 图片

4)删掉策略root@Ceph01 ~# s3cmd delpolicy s3://kpyun-buckets3://kpyun-bucket/: Policy deletedroot@Ceph01 ~# curl -s http://10.0.0.3/kpyun-bucket/s.png<Code>AccessDenied</Code><Message>missing s3:GetObject permission</Message>'访问被拒绝,因为缺少 s3:GetObject 权限'root@Ceph01 ~# s3cmd info s3://kpyun-bucket | grep Policy Policy: none
5)应用到 kpyun-backup 桶root@Ceph01 ~# sed -i s#kpyun-bucket#kpyun-backup# /root/anon-policy.jsonroot@Ceph01 ~# s3cmd setpolicy /root/anon-policy.json s3://kpyun-backups3://kpyun-backup/: Policy updatedroot@Ceph01 ~# curl -s http://10.0.0.3/kpyun-backup/s3test.txthello kpyun ceph storage✅️ 不用 secret_key 也能读到
6)删除桶 + 桶中的文件root@Ceph01 ~# s3cmd ls s3://kpyun-backup | wc -l5root@Ceph01 ~# s3cmd rb -r s3://kpyun-backup-r = --recursive # 递归操作(处理子目录)WARNING: Bucket is not empty. Removing all the objects from it first. This may take some time...delete: 's3://kpyun-backup/auth.log'delete: 's3://kpyun-backup/auth.log.1'delete: 's3://kpyun-backup/auth.log.2.gz'delete: 's3://kpyun-backup/big.bin'delete: 's3://kpyun-backup/s3test.txt'Bucket 's3://kpyun-backup/' removed`彻底删除整个桶(包括里面的所有文件)`Python 操作对象存储
1)准备隔离环境jiuzhao@Ubuntu ~$ pipx --version1.8.0jiuzhao@Ubuntu ~$ uv --versionuv 0.11.29 (x86_64-unknown-linux-gnu)jiuzhao@Ubuntu ~$ uv venv test --python 3.14Using CPython 3.14.4 interpreter at: /usr/bin/python3.14Creating virtual environment at: testActivate with: source test/bin/activatejiuzhao@Ubuntu ~$ source ./test/bin/activate(test) jiuzhao@Ubuntu ~$ uv pip install -i https://pypi.tuna.tsinghua.edu.cn/simple boto3# 用清华镜像
2)写操作脚本(test) jiuzhao@Ubuntu ~$ cat > ./test/rgw-kpyun.py <<'EOF'# 导入boto3库,这是AWS官方Python SDK,用于操作S3对象存储import boto3# 从botocore导入Config类,用于配置客户端参数from botocore.client import Config
# 设置访问密钥(Access Key和Secret Key)# 这些相当于用户名和密码,用于身份认证access_key = "42GA8G0I7JGTVGX43NPK"secret_key = "8pJrGe843L6C8xemXAi3RhA2OE28WIguFcjKg247"
# 创建S3客户端对象,用于后续操作# 这个客户端相当于一个"连接器",连接到S3服务s3 = boto3.client("s3", # endpoint_url指定S3服务的访问地址(这里是本地的RGW服务) endpoint_url="http://10.0.0.3", # 传入访问密钥 aws_access_key_id=access_key, aws_secret_access_key=secret_key, # 指定使用签名版本4(S3v4),这是较新的认证方式 config=Config(signature_version="s3v4"),)
# 创建一个名为"kpyun-rgw"的存储桶(Bucket)# 存储桶类似于文件系统中的根目录,用来存放对象s3.create_bucket(Bucket="kpyun-rgw")
# 列出所有存储桶并打印出来print("=== buckets ===")# list_buckets()返回所有存储桶列表,["Buckets"]取出这个列表for b in s3.list_buckets()["Buckets"]: # 打印每个存储桶的名称和创建时间 print("{name}\t{created}".format(name=b["Name"], created=b["CreationDate"]))
# 在"kpyun-rgw"存储桶中上传一个对象(文件)# put_object相当于上传文件,Key是文件名,Body是文件内容s3.put_object(Bucket="kpyun-rgw", Key="test.txt", Body=b"Hello World")
# 列出"kpyun-rgw"存储桶中的所有对象并打印print("=== objects in kpyun-rgw ===")# list_objects_v2列出指定存储桶中的所有对象,["Contents"]取出对象列表for o in s3.list_objects_v2(Bucket="kpyun-rgw")["Contents"]: # 打印每个对象的文件名和大小 print("{name}\t{size}".format(name=o["Key"], size=o["Size"]))
# 从存储桶中下载"test.txt"对象# get_object相当于下载文件obj = s3.get_object(Bucket="kpyun-rgw", Key="test.txt")print("=== get test.txt ===")# 读取对象内容(Body是文件内容),然后decode()将字节转为字符串打印print(obj["Body"].read().decode())EOF
3)运行(test) jiuzhao@Ubuntu ~$ python ./test/rgw-kpyun.py=== buckets ===kpyun-bucket 2026-08-29 02:21:44.625000+00:00kpyun-rgw 2026-08-29 10:38:04.612000+00:00=== objects in kpyun-rgw ===test.txt 11=== get test.txt ===Hello World✅️ 建桶、写对象、读对象一条龙通了
脚本执行条件: 1)只需要 Python环境 + boto3库 2)网络能连通 Ceph RGW 服务地址(http://10.0.0.3) 3)有正确的访问密钥(代码里已写好)想让外部匿名访问这个 test.txt,同样先给 kpyun-rgw 加匿名读策略再 curl:
root@Ceph01 ~# sed -i s#kpyun-backup#kpyun-rgw# /root/anon-policy.jsonroot@Ceph01 ~# s3cmd setpolicy /root/anon-policy.json s3://kpyun-rgws3://kpyun-rgw/: Policy updatedroot@Ceph01 ~# curl -s http://10.0.0.3/kpyun-rgw/test.txtHello World'匿名读也 OK,最后记得 delpolicy 收尾'root@Ceph01 ~# s3cmd delpolicy s3://kpyun-rgwroot@Ceph01 ~# curl -s http://10.0.0.3/kpyun-rgw/test.txt<Code>AccessDenied</Code><Message>missing s3:GetObject permission</Message>Swift 操作对象存储
概念铺垫(Swift 数据模型 Account -> container -> Object、subuser 权限)已在 网关概述 讲完,这里不重复
- 本章客户端:
python-swiftclient,提供swift命令 + Python API - Swift 的
subuser挂在前面创建的jiuzhao用户下,登录用jiuzhao:swift+ 独立的swift_key - Swift 的
container等价于 S3 的bucket(同一命名空间,下面验证)
📌 S3 与 Swift 是同一命名空间:RGW 同一套后端同时服务两种协议,所以 jiuzhao 用户下的桶,s3cmd 和 swift 都能看到 —— 末尾 swift list 直接验证
Swift 命令行实战
1)创建子用户(挂在 jiuzhao 这个 user 下)root@Ceph01 ~# radosgw-admin subuser create --uid=jiuzhao --subuser=jiuzhao:swift --access=full`子用户不指定 --access 默认是「无权限」(empty),必须显式指定`# --access=full 是最高档,子用户能建容器、上传、删除全覆盖(读 + 写 + 管理){ ... "swift_keys": [ { "user": "jiuzhao:swift", "secret_key": "cBVi4iEZ8SG6XdVpKm6uQ9GljgsXMxFxhxj7e1qH", "active": true, "create_date": "2026-08-29T11:21:38.256534Z" } ]}'拿到 swift 子用户的 secret_key'
2)装 swift 客户端(test) jiuzhao@Ubuntu ~$ uv pip install -i https://pypi.tuna.tsinghua.edu.cn/simple python-swiftclient==4.6.0(test) jiuzhao@Ubuntu ~$ swift --versionpython-swiftclient 4.6.0
3)建容器 + 上传文件(认证端点用 rgw 的 /auth)(test) jiuzhao@Ubuntu ~$ swift -A http://10.0.0.3/auth -U jiuzhao:swift -K cBVi4iEZ8SG6XdVpKm6uQ9GljgsXMxFxhxj7e1qH post kpyun-swift-A http://10.0.0.3/auth # Auth URL 认证服务的地址,告诉swift去哪里验证身份-U jiuzhao:swift # User 用户名,格式是 账号:用户名,这里是 jiuzhao 账号下的 swift 用户-K cBVi4iEZ8SG6Xd... # Key 用户的认证密钥(相当于密码)post # 操作命令,创建或更新容器(这里是创建)kpyun-swift # 容器名称,要创建的容器名字(test) jiuzhao@Ubuntu ~$ swift -A http://10.0.0.3/auth -U jiuzhao:swift -K cBVi4iEZ8SG6XdVpKm6uQ9GljgsXMxFxhxj7e1qH upload kpyun-swift /etc/os-release /etc/hosts /etc/fstab# 用同样的认证信息登录,然后上传三个文件到 kpyun-swift 容器里etc/os-releaseetc/hostsetc/fstab
4)查看容器内容(test) jiuzhao@Ubuntu ~$ swift -A http://10.0.0.3/auth -U jiuzhao:swift -K cBVi4iEZ8SG6XdVpKm6uQ9GljgsXMxFxhxj7e1qH list kpyun-swiftetc/fstabetc/hostsetc/os-release✅️ 三个文件进容器了
5)用环境变量简化(避免每次敲 -A/-U/-K)(test) jiuzhao@Ubuntu ~$ cat > ./test/.swift <<'EOF'export ST_AUTH=http://10.0.0.3/authexport ST_USER=jiuzhao:swiftexport ST_KEY=cBVi4iEZ8SG6XdVpKm6uQ9GljgsXMxFxhxj7e1qHEOF(test) jiuzhao@Ubuntu ~$ source ./test/.swift(test) jiuzhao@Ubuntu ~$ swift listkpyun-bucketkpyun-rgwkpyun-swift🌿 重点看:swift list 把前面 s3cmd 建的 bucket 也列出来了!'证明 S3 和 Swift 在 RGW 里共用同一套后端,并不是两套隔离的桶'(test) jiuzhao@Ubuntu ~$ swift list kpyun-bucket --lh--lh # 这个参数就是专门用来显示人类可读格式的(显示详细信息:大小、时间、类型) 40M 2026-08-29 03:05:33 application/octet-stream big.bin2.0M 2026-08-29 02:37:28 video/webm bilibili.webm568K 2026-08-29 02:37:28 image/png s.png 42MPrometheus监控Ceph集群
root@Ceph01 ~# ceph orch lsNAME PORTS RUNNING REFRESHED AGE PLACEMENTalertmanager ?:9093,9094 1/1 8m ago 5w count:1ceph-exporter ?:9926 3/3 8m ago 5w *grafana ?:3000 1/1 8m ago 5w count:1node-exporter ?:9100 3/3 8m ago 5w *prometheus ?:9095 1/1 8m ago 5w count:1`Ceph 集群已经通过 cephadm 成功部署并运行了完整的监控栈`
root@Ceph01 ~# ceph orch ps --service_name prometheusNAME HOST PORTS STATUS REFRESHED AGEprometheus.Ceph01 Ceph01 *:9095 running (34m) 3m ago 5w# 运行在Ceph01的9095端口
http://10.0.0.3:9095/targets
root@Ceph01 ~# ceph orch ps --service_name grafanaNAME HOST PORTS STATUS REFRESHED AGEgrafana.Ceph01 Ceph01 *:3000 running (37m) 6m ago 5w
https://10.0.0.3:3000/
root@Ceph01 ~# ceph orch ps --service_name alertmanagerNAME HOST PORTS STATUS REFRESHED AGEalertmanager.Ceph01 Ceph01 *:9093,9094 running (52m) 71s ago 5w# 9093 负责对外提供服务,9094 负责内部集群通信
http://10.0.0.3:9093/#/status
root@Ceph01 ~# docker exec -it ceph-a9b741cd-856a-11f1-b216-525400c6b9fa-alertmanager-Ceph01 sh/alertmanager $ whoaminobody/alertmanager $ ps -efPID USER TIME COMMAND 1 nobody 0:00 /sbin/docker-init -- /bin/alertmanager --web.listen-address=:9093 --cluster.listen-address=:9094 --config.file=/etc/alertmanager/alertmanager.yml 8 nobody 0:02 /bin/alertmanager --web.listen-address=:9093 --cluster.listen-address=:9094 --config.file=/etc/alertmanager/alertmanager.yml 101 nobody 0:00 sh 108 nobody 0:00 ps -ef/alertmanager $ ls -lh /etc/alertmanager/total 8K-rw------- 1 nobody nobody 588 Jul 22 02:09 alertmanager.ymldrwxr-xr-x 2 nobody nobody 4.0K Jul 22 01:15 dataCeph集群性能压测
Ceph压测核心应用场景:
- 初建基线:集群上线时跑一次,记录顺序吞吐(MB/s)和随机IOPS,作为后续所有对比的“标尺”
- 调试验证:改完内核/配置参数后对比重测,用数据判断优化是否有效,避免凭感觉调优
- 问题定界:用户反馈读/写慢时,分方向压测(区分读/写),结合
ceph -w看集群状态,快速判断是磁盘瓶颈、网络问题还是集群自愈任务占资源
顺带一句:压测前务必确认集群没有正在进行的Recovery/Backfill,否则数据失真
| 压测工具 | 测什么 | 关键命令 | 运行位置 |
|---|---|---|---|
rados bench | 对象/存储池(RADOS) | rados bench -p <pool> <sec> write/seq/rand | 集群节点 |
rbd bench | 块设备(RBD) | rbd bench --io-type write/read <image> --pool=<pool> | 集群节点 |
fio | 文件系统(CephFS) | fio --rw=randwrite --direct=1 --ioengine=libaio ... | CephFS 挂载客户端 |
s3cmd | 对象网关(RGW) | time s3cmd put/get ... | 集群节点 |
rados bench对象池压测
📌 一句话:rados bench 是 Ceph 自带的性能测试工具,专门评估 RADOS 存储池的写、顺序读、随机读性能
参考链接: https://www.ibm.com/docs/zh/storage-ceph/7.0.0?topic=benchmark-benchmarking-ceph-performance
rados bench 选项速查表(本机实测命令逐项拆解):
| 选项 | 本机取值 | 作用 |
|---|---|---|
-p | bench | 指定被压测的存储池 |
30 | 30 秒 | 压测时长(位置参数,单位秒) |
write | 写入 | 压测模式:write 顺序写 / seq 顺序读 / rand 随机读 |
-b | 4M | 单次写入的块大小(默认 4M) |
-t | 16 | 并发线程数,全程维持 16 个在途写(默认 16) |
--no-cleanup | 开启 | 测完不删测试对象,留给后面 seq/rand 读测试复用 |
1)环境准备:建一个压测池root@Ceph01 ~# ceph osd pool create benchpool 'bench' created
2)写性能测试(4M 块 × 16 并发 × 30 秒)root@Ceph01 ~# rados bench -p bench 30 write -b 4M -t 16 --no-cleanuphints = 1# 向 OSD 发送分配提示(hints):1 = 开启,新版默认行为Maintaining 16 concurrent writes of 4194304 bytes to objects of size 4194304 for up to 30 seconds or 0 objectsObject prefix: benchmark_data_Ceph01_23846# 测试对象统一用这个前缀命名,--no-cleanup 留下的数据好认好清 sec Cur ops started finished avg MB/s cur MB/s last lat(s) avg lat(s) 0 16 16 0 0 0 - 0 1 16 50 34 135.946 136 0.293907 0.40523 2 16 116 100 199.948 264 0.0761148 0.29383 3 16 186 170 226.613 280 0.0895038 0.267447# 第 3 秒瞬时带宽冲上 280 MB/s → 汇总里的 Max bandwidth 就是它# 前几秒 200+ 是缓存红利,后面掉到 130 左右才是真实持续速率... 9 16 443 427 189.662 16 1.31546 0.279643# 瞬时值也会猛掉到 16,虚拟机环境抖动是常态... 19 16 728 712 149.838 28 1.31587 0.4042382026-08-31T16:09:09.500209+0800 min lat: 0.0345262 max lat: 2.76964 avg lat: 0.423316# 中途打一行带时间戳的延迟快照(min/max/avg 三件套) sec Cur ops started finished avg MB/s cur MB/s last lat(s) avg lat(s) 20 16 738 722 144.346 40 2.13048 0.423316# 这一秒出现 2.13s 的慢 IO,长尾延迟的苗头... 28 16 919 903 128.959 12 1.07598 0.479168 29 16 922 906 124.925 12 1.99916 0.48361# 第 28/29 秒掉到 12 MB/s → 汇总里的 Min bandwidth... 30 16 960 944 125.827 152 0.204938 0.506318Total time run: 30.1391Total writes made: 960Write size: 4194304Object size: 4194304Bandwidth (MB/sec): 127.409 # 带宽 *****Stddev Bandwidth: 92.6885Max bandwidth (MB/sec): 280Min bandwidth (MB/sec): 12Average IOPS: 31 # IOPS *****Stddev IOPS: 23.1721Max IOPS: 70Min IOPS: 3Average Latency(s): 0.501605 # 平均延迟 *****Stddev Latency(s): 0.582611Max latency(s): 2.76964Min latency(s): 0.0345262实时监控字段说明(上面每秒刷新的 8 列):
| 字段 | 含义 | 本机观察 |
|---|---|---|
sec | 开跑后的秒数 | 0 → 30 |
Cur ops | 此刻在途 IO 数 | 恒为 16,并发打满 ✅️ |
started / finished | 累计发起 / 完成的 IO 数 | 30s 时 960 / 944,差值 = 还在飞的 16 个 |
avg MB/s | 开跑至今的平均带宽 | 从 240 一路缓降到 126,评估看这列 |
cur MB/s | 最近 1 秒的瞬时带宽 | 12~280 之间剧烈跳 |
last lat(s) | 最近一次完成 IO 的延迟 | 0.05 ~ 2.13 |
avg lat(s) | 平均延迟 | 从 0.26 爬到 0.51 |
压测结果汇总解读(末尾汇总才是重点):
| 结果字段 | 实测值 | 说明 |
|---|---|---|
Total time run | 30.1391 | 实际总时长,略超 30s 是因为在途 IO 要收尾 |
Total writes made | 960 | 完成的写次数,每次写满一个 4M 对象 |
Write size / Object size | 4194304 | 单次写 4M;一个对象一次写满 |
Bandwidth (MB/sec) | 127.409 | 平均带宽 = 960×4M ÷ 30.14s,核心指标 ***** |
Stddev Bandwidth | 92.6885 | 秒级带宽标准差,达均值 7 成 → 波动剧烈 |
Max / Min bandwidth | 280 / 12 | 瞬时峰值(第 3 秒)/ 谷值(第 28-29 秒) |
Average IOPS | 31 | 每秒平均完成 31 个 4M 对象写,核心指标 ***** |
Stddev / Max / Min IOPS | 23.17 / 70 / 3 | IOPS 的波动范围 |
Average Latency(s) | 0.501605 | 单次 4M 写平均耗时 0.50s,核心指标 ***** |
Stddev Latency(s) | 0.582611 | 延迟标准差比均值还大 → 快慢差距悬殊 |
Max / Min latency(s) | 2.76964 / 0.0345262 | 最慢一次 2.77s(长尾延迟)/ 最快 35ms |
💡 这份结果说明了什么:
- 三个核心指标自洽:带宽 = 写入总量 ÷ 时长(960×4M÷30.14s≈127),IOPS = 并发 ÷ 平均延迟(16÷0.50≈31),两条公式都验算通过 ✅️
- 127 MB/s 是客户端视角的吞吐,3 副本集群底层实际落盘 ≈ ×3 ≈ 382 MB/s
- Stddev 92.7(均值 127):虚拟机磁盘抖动的正常表现,容量评估看 avg,别被瞬时值带偏
- Max latency 2.77s:典型的长尾延迟(副本对齐 + 慢 IO),延迟敏感业务要盯它而不是均值
'跑完用 bc 把关键指标验算一遍'root@Ceph01 ~# echo 4194304/1024/1024|bc4# 4194304 bytes = 4Mroot@Ceph01 ~# echo 960*4/1024|bc3# 960 次 × 4M = 3.75G,客户端视角的写入总量root@Ceph01 ~# echo 960*4/30.1391|bc127# 带宽 = 写入总量 ÷ 时长:3840M ÷ 30.14s ≈ 127 MB/s,对上了 ✅️root@Ceph01 ~# echo 16/0.501605|bc31# IOPS = 并发数 ÷ 平均延迟(Little's Law):16 ÷ 0.50 ≈ 31,又对上了 ✅️
'再看看落盘:3 副本到底占了多少空间'root@Ceph01 ~# rados df | egrep "bench|POOL_NAME"POOL_NAME USED OBJECTS CLONES COPIES ...bench 11 GiB 961 0 2883 ...# 960 个数据对象 + 1 个 bench 元数据对象(benchmark_last_metadata)# COPIES = 961 × 3 = 2883 → 3 副本实锤root@Ceph01 ~# echo 2883/961|bc3root@Ceph01 ~# echo 960*4*3/1024|bc11# 3.75G × 3 副本 ≈ 11.25G ≈ USED 的 11 GiB,空间账也对上了 ✅️
3)顺序读测试(每个对象从头到尾读一遍)'⚠️ 读测试必须紧跟写测试——seq/rand 靠 bench 元数据定位数据,中途 cleanup 过就得先重写一遍,否则报 Must write data'root@Ceph01 ~# rados bench -p bench 10 seq -t 16 --no-cleanuphints = 1 sec Cur ops started finished avg MB/s cur MB/s last lat(s) avg lat(s) 0 1 1 0 0 0 - 0 1 16 232 216 862.656 864 0.0874515 0.0693965 2 16 449 433 865.244 868 0.0434486 0.0704194 3 16 662 646 857.177 852 0.0712364 0.0714316 4 16 850 834 830.928 752 0.0755911 0.073763Total time run: 4.62293Total reads made: 960Read size: 4194304Object size: 4194304Bandwidth (MB/sec): 830.642 # 顺序读带宽,写入的 6.5 倍Average IOPS: 207Stddev IOPS: 13.772Max IOPS: 217Min IOPS: 188Average Latency(s): 0.0747287Max latency(s): 0.184588Min latency(s): 0.010089# 4.6s 就结束:960 个对象各读一遍(正好是刚写入的 960 个,一个不差),读完即止不凑满 10s
4)随机读测试(先清缓存,逼数据从磁盘走)root@Ceph01 ~# echo 3 | sudo tee /proc/sys/vm/drop_caches && sudo sync3root@Ceph01 ~# rados bench -p bench 10 rand -t 16 --no-cleanuphints = 1 sec Cur ops started finished avg MB/s cur MB/s last lat(s) avg lat(s) 0 0 0 0 0 0 - 0 1 16 250 234 934.614 936 0.0481291 0.0647051 2 16 496 480 959.052 984 0.0239671 0.0631288... 9 16 2253 2237 989.17 1000 0.0592999 0.0622191Total time run: 10.0415Total reads made: 2487# 2487 是读次数不是对象数——对象还是那 960 个,随机抽着平均每个被读 ~2.6 次Read size: 4194304Object size: 4194304Bandwidth (MB/sec): 990.691 # 接近 1GB/s,比顺序读还高Average IOPS: 247Stddev IOPS: 8.69067Max IOPS: 257Min IOPS: 234Average Latency(s): 0.0623715Max latency(s): 0.200771Min latency(s): 0.00440801💡 为什么随机读前要清缓存:
echo 3:1只清 pagecache,2只清 dentries/inodes,3 = 两者都清sudo tee而不是>:>重定向由 shell 先执行,sudo 管不到它;tee在提权后写入,才写得进/procsync:先把脏页刷回磁盘再清,清得干净- 不清的后果:上一步刚读过的对象还在内存里,随机读全命中缓存,测的是内存不是集群
- ⚠️ 生产环境别乱清——缓存一掉,业务 IO 瞬间洪峰
4M 块读写对比(写 30s / 读 10s,同为 16 并发):
| 模式 | 带宽 MB/s | IOPS | 平均延迟 |
|---|---|---|---|
| 4M 写 | 127.4 | 31 | 0.50s |
| 4M 顺序读 | 830.64 | 207 | 0.075s |
| 4M 随机读 | 990.69 | 247 | 0.062s |
💡 读写对比结论:
- 读带宽是写的 6~8 倍:写要 3 副本全部落盘才算完成,读随便挑一份副本就行
- 随机读比顺序读还高:960 个对象打散在 12 个 OSD 上,随机抽反而把负载摊得更匀
- 读延迟 0.06
0.07s vs 写 0.50s,差距同样是 68 倍
5)清理测试数据root@Ceph01 ~# rados df | egrep "bench|POOL_NAME"POOL_NAME USED OBJECTS CLONES COPIES ...bench 11 GiB 961 0 2883 ...# 960 个数据对象 + 1 个 bench 元数据对象;11 GiB = 960 × 4M × 3 副本root@Ceph01 ~# rados -p bench cleanupRemoved 960 objectsroot@Ceph01 ~# rados df | egrep "bench|POOL_NAME"bench 0 B 0 0 0 ...✅️ 压测数据清空,bench 池可以留作后用rbd bench块设备压测
📌 一句话:rbd bench 测试对 RBD 块设备的写入/读取性能,缺省 4K 块 × 16 线程 × 1G 总量 × 顺序模式——本次把默认值全部显式写出,顺便记住它们
参考链接: https://www.ibm.com/docs/zh/storage-ceph/7.0.0?topic=benchmark-benchmarking-ceph-block-performance
rbd bench 选项速查表(⭐ = 默认值):
| 选项 | 本机取值 | 作用 |
|---|---|---|
kpyun --pool bench | bench/kpyun | 被压测的镜像,也可直接写成 bench/kpyun 一个参数 |
--io-type | write / read | IO 类型:写 / 读(还有 readwrite 混合) |
--io-size | 4K ⭐ | 单次 IO 的块大小,默认 4K |
--io-threads | 16 ⭐ | 并发线程数,默认 16 |
--io-total | 1G ⭐ | 总 IO 量,写/读满 1G 即止,默认 1G |
--io-pattern | seq ⭐ | IO 模式:默认顺序 seq,可选随机 rand |
1)创建测试块设备(10G 精简配置镜像)root@Ceph01 ~# rbd ls -l -p bench# 空root@Ceph01 ~# rbd create --size 10G bench/kpyunroot@Ceph01 ~# rbd ls -l -p benchNAME SIZE PARENT FMT PROT LOCKkpyun 10 GiB 2
2)块设备写性能测试(4K × 16 线程 × 1G 顺序写,全是默认值,显式写出便于对照)root@Ceph01 ~# rbd bench --io-type write --io-size 4K --io-threads 16 --io-total 1G --io-pattern seq kpyun --pool benchbench type write io_size 4096 io_threads 16 bytes 1073741824 pattern sequential# 头行回显全部参数:4096 字节块 / 16 线程 / 共 1G / 顺序模式 SEC OPS OPS/SEC BYTES/SEC 1 25152 25218.2 99 MiB/s 2 57392 28732.5 112 MiB/s 3 92640 30885.1 121 MiB/s... 7 220000 32501.9 127 MiB/s# 爬到 127 MiB/s 后稳定...elapsed: 8 ops: 262144 ops/sec: 30145.1 bytes/sec: 118 MiB/s# elapsed 汇总行:8 秒写完 262144 笔,平均 30145 IOPS / 118 MiB/s 平均带宽(吞吐量)'1G 数据 8 秒写完,小 4K 块 IOPS 轻松上 3 万'实时字段说明(每秒刷新的 4 列 + 汇总行):
| 字段 | 含义 |
|---|---|
SEC | 开跑后的秒数 |
OPS | 累计完成的 IO 笔数 |
OPS/SEC | 最近 1 秒的瞬时 IOPS |
BYTES/SEC | 最近 1 秒的瞬时带宽 |
elapsed 汇总行 | 总时长 / 总笔数 / 平均 IOPS / 平均带宽,最终看这行 |
'照例用 bc 验算一遍'root@Ceph01 ~# echo 1073741824/4096|bc262144# 1G ÷ 4K = 262144 笔,和 elapsed 里的 ops 分毫不差root@Ceph01 ~# echo 30145*4096/1024/1024|bc117# 带宽 = IOPS × 块大小:30145 × 4K ≈ 118 MiB/s,对上了 ✅️root@Ceph01 ~# echo 5832*4096/1024/1024|bc22# 下面读性能测试同样成立:5832 × 4K ≈ 23 MiB/s ✅️3)清缓存后块设备读性能测试root@Ceph01 ~# echo 3 | sudo tee /proc/sys/vm/drop_caches && sudo sync3# 刚写完的数据还挂在内存里,不清的话读测试全是缓存命中,分数虚高root@Ceph01 ~# rbd bench --io-type read --io-size 4K --io-threads 16 --io-total 1G --io-pattern seq kpyun --pool benchbench type read io_size 4096 io_threads 16 bytes 1073741824 pattern sequential SEC OPS OPS/SEC BYTES/SEC 1 5856 5919.31 23 MiB/s 2 11200 5627.65 22 MiB/s... 11 59440 4827.8 19 MiB/s# 中段一度掉到 19 MiB/s... 18 105136 6711.69 26 MiB/s# 全程峰值也就 26 MiB/s... 44 256048 5973.16 23 MiB/selapsed: 44 ops: 262144 ops/sec: 5832.38 bytes/sec: 23 MiB/s# 读同样的 1G 花了 44 秒,是写入时长的 5.5 倍
4)清理测试镜像root@Ceph01 ~# rbd remove bench/kpyunroot@Ceph01 ~# rbd ls -l -p bench# 空✅️ 块设备压测结束,环境还原rbd 4K 读写对比(同一镜像、同样 1G 总量):
| 模式 | 耗时 | IOPS | 带宽 |
|---|---|---|---|
| 4K 顺序写 | 8s | 30145 | 118 MiB/s |
| 4K 顺序读(清缓存后) | 44s | 5832 | 23 MiB/s |
💡 写比读快 5 倍?别怀疑,真复现了,原因在这:
- 写:OSD 先把数据塞进内存缓冲/流水线就给你回话,盘还没落完这单就算完成了——先应答后落盘,所以快得飞起
- 读:清完缓存等于裸奔,每一笔 4K 都得实打实去虚拟磁盘里刨回来,小 IO 最怕来回折腾,延迟全被放大,自然就慢了
- 生产上 RBD 读得快不快,全靠 OSD 内存缓存 + SSD 撑着,冷数据第一次读就是这么拉,急也急不来
FIO文件系统压测
📌 一句话:用 fio 工具对 CephFS 文件系统做基准测试,测 4k 块在不同 iodepth 下的随机写性能,找到饱和点
参考链接: https://www.ibm.com/docs/zh/storage-ceph/7.0.0?topic=benchmark-benchmarking-cephfs-performance
1)安装 fiojiuzhao@Ubuntu ~$ apt -y install fio
2)服务端创建 cephfs 用户并导出 keyroot@Ceph01 ~# ceph mds statops-cephfs:1 {0=ops-cephfs.Ceph02.bwbckx=up:active} 1 up:standbyroot@Ceph01 ~# ceph auth add client.cephfs mon 'allow r' mds 'allow rw' osd 'allow rwx'root@Ceph01 ~# ceph auth get client.cephfs[client.cephfs] key = AQA5wGRqh4c+IRAAw4a9Nh3Ukek0A4ziYVgREg== caps mds = "allow rw" caps mon = "allow r" caps osd = "allow rwx"
3)客户端挂载 cephfsjiuzhao@Ubuntu ~$ mkdir -p /home/jiuzhao/testjiuzhao@Ubuntu ~$ sudo mount -t ceph 10.0.0.3:6789,10.0.0.4:6789,10.0.0.5:6789:/ /home/jiuzhao/test -o name=cephfs,secret=AQA5wGRqh4c+IRAAw4a9Nh3Ukek0A4ziYVgREg==# 虽有警告,无伤大雅jiuzhao@Ubuntu ~$ df -h | grep /home/jiuzhao/test10.0.0.3:6789,10.0.0.4:6789,10.0.0.5:6789:/ 57G 0 57G 0% /home/jiuzhao/testFIO 核心参数速查表:
| 参数 | 取值 | 含义 |
|---|---|---|
--name=randwrite | randwrite | 任务名称(自定义,方便区分多个任务) |
--rw=randwrite | randwrite | 随机写模式(测试磁盘随机写入性能) |
--direct=1 | 1 | 开启裸 IO / 直写,绕过 pagecache,测真实磁盘性能 |
--ioengine=libaio | libaio | 使用 Linux 原生异步 IO 引擎(高性能标准引擎) |
--bs=4k | 4k | 单次 IO 块大小为 4KB(数据库/业务最常用块大小) |
--iodepth=1 | 1 | IO 队列深度 = 1(单队列,测单线程延迟) |
--size=5G | 5G | 测试文件总大小 5GB |
--runtime=60 | 60 | 测试时长 60 秒 |
--group_reporting=1 | 1 | 汇总所有线程统一输出报告(方便查看) |
4)4k 随机写,压不同 iodepth(1 / 32 / 64 / 128),看并行度与饱和点jiuzhao@Ubuntu ~$ cd /home/jiuzhao/testjiuzhao@Ubuntu test$ fio --name=randwrite1 --rw=randwrite --direct=1 --ioengine=libaio --bs=4k --iodepth=1 --size=5G --runtime=60 --group_reporting=1... write: IOPS=274, BW=1097KiB/s (1123kB/s)(64.3MiB/60002msec); 0 zone resets...Run status group 0 (all jobs): WRITE: bw=1097KiB/s (1123kB/s), io=64.3MiB, run=60002-60002msec# iodepth=1: 单队列,延迟最低(3.6ms)但 IOPS 最低,60s 只写了 64M
jiuzhao@Ubuntu test$ fio --name=randwrite32 --rw=randwrite --direct=1 --ioengine=libaio --bs=4k --iodepth=32 --size=5G --runtime=60 --group_reporting=1... write: IOPS=1453, BW=5815KiB/s (5955kB/s)(341MiB/60010msec); 0 zone resets... WRITE: bw=5815KiB/s (5955kB/s), io=341MiB, run=60010-60010msec# iodepth=32: 并行度上来,IOPS 涨了 5 倍,延迟 22ms
jiuzhao@Ubuntu test$ fio --name=randwrite64 --rw=randwrite --direct=1 --ioengine=libaio --bs=4k --iodepth=64 --size=5G --runtime=60 --group_reporting=1... write: IOPS=1640, BW=6561KiB/s (6718kB/s)(385MiB/60017msec); 0 zone resets... WRITE: bw=6561KiB/s (6718kB/s), io=385MiB, run=60017-60017msec# iodepth=64: 甜蜜点,IOPS 和带宽爬到最高,延迟 39ms
jiuzhao@Ubuntu test$ fio --name=randwrite128 --rw=randwrite --direct=1 --ioengine=libaio --bs=4k --iodepth=128 --size=5G --runtime=60 --group_reporting=1... write: IOPS=656, BW=2625KiB/s (2688kB/s)(154MiB/60040msec); 0 zone resets... WRITE: bw=2625KiB/s (2688kB/s), io=154MiB, run=60040-60040msec# iodepth=128: 队列过深,饱和了!IOPS 暴跌到 656,延迟爆炸到 195ms
jiuzhao@Ubuntu test$ ll -h | grep randwrite-rw-r--r-- 1 jiuzhao jiuzhao 5.0G Aug 31 18:45 randwrite1.0.0-rw-r--r-- 1 jiuzhao jiuzhao 5.0G Aug 31 18:46 randwrite32.0.0-rw-r--r-- 1 jiuzhao jiuzhao 5.0G Aug 31 18:47 randwrite64.0.0-rw-r--r-- 1 jiuzhao jiuzhao 5.0G Aug 31 18:48 randwrite128.0.0# 每个任务各生成了一个 5G 的测试文件(跑完已 rm 清理掉)4k 随机写结果对比:
| iodepth | IOPS | 带宽 | 60s 写入量 | 平均延迟 |
|---|---|---|---|---|
| 1 | 274 | 1097 KiB/s | 64 MiB | 3.6 ms |
| 32 | 1453 | 5815 KiB/s | 341 MiB | 22 ms |
| 64 | 1640 | 6561 KiB/s | 385 MiB | 39 ms |
| 128 | 656 ↓ | 2625 KiB/s ↓ | 154 MiB | 195 ms ↑ |
💡 结论:iodepth 不是越大越好,过了甜蜜点就饱和
- 1→32→64:并行度上来,IOPS 和带宽一路爬,甜蜜点在 iodepth≈64(1640 IOPS / 6.5 MiB/s)
- 128:队列过深,后端处理不过来,IOPS 从 1640 暴跌到 656,延迟从 39ms 爆炸到 195ms
- 生产环境要根据这套曲线找到拐点,盲目加大 iodepth 适得其反
RGW对象网关压测
📌 一句话:用 s3cmd + time 命令测量 RGW 对象存储网关的上传、下载、列表耗时,反推读写带宽
1)确认桶里的现有对象root@Ceph01 ~# s3cmd ls s3://kpyun-bucket2026-08-29 03:05 41943040 s3://kpyun-bucket/big.bin2026-08-29 02:37 2053779 s3://kpyun-bucket/bilibili.webm2026-08-29 02:37 582512 s3://kpyun-bucket/s.png
2)生成 1G 测试文件并上传测量耗时root@Ceph01 ~# dd if=/dev/zero of=/root/bigfile01 bs=1M count=10241024+0 records in1024+0 records out1073741824 bytes (1.1 GB, 1.0 GiB) copied, 2.3 s, 467 MB/sroot@Ceph01 ~# ls -lh /root/bigfile01-rw-r--r-- 1 root root 1.0G Aug 31 19:42 bigfile01root@Ceph01 ~# time s3cmd put /root/bigfile01 s3://kpyun-bucketupload: '/root/bigfile01' -> 's3://kpyun-bucket/bigfile01' [part 1 of 69, 15MB] [1 of 1] 15728640 of 15728640 100% in 0s 77.00 MB/s done......upload: '/root/bigfile01' -> 's3://kpyun-bucket/bigfile01' [part 69 of 69, 4MB] [1 of 1] 4194304 of 4194304 100% in 0s 62.68 MB/s donereal 0m15.312s # 墙钟 15.3s,这是算带宽用的时间user 0m4.183s # 用户态 CPU 时间sys 0m2.151s # 内核态 CPU 时间root@Ceph01 ~# echo 1073741824/15.3/1024/1024|bc66'1G / 15.3s ≈ 66MB/s,rgw 上传性能就这么算出来'
3)下载文件并测量耗时root@Ceph01 ~# time s3cmd get s3://kpyun-bucket/bigfile01 /tmp/download: 's3://kpyun-bucket/bigfile01' -> '/tmp/bigfile01' (1073741824 bytes in 4.6 seconds, 220.53 MB/s) [1 of 1]real 0m4.918s # 墙钟不到 5 秒user 0m1.703ssys 0m1.424sroot@Ceph01 ~# echo 1073741824/4.9/1024/1024|bc209'下载 ≈ 209MB/s,下载比上传快 3 倍多'
4)列出桶中所有对象并测量响应时间root@Ceph01 ~# time s3cmd ls s3://kpyun-bucket/2026-08-29 03:05 41943040 s3://kpyun-bucket/big.bin2026-08-31 11:42 1073741824 s3://kpyun-bucket/bigfile012026-08-29 02:37 2053779 s3://kpyun-bucket/bilibili.webm2026-08-29 02:37 582512 s3://kpyun-bucket/s.png
real 0m0.249s # 250 毫秒,毫秒级响应user 0m0.119ssys 0m0.036s'压测完记得清掉测试对象:s3cmd del s3://kpyun-bucket/bigfile01'RGW 对象网关压测结果(1G 单文件,直传未分片):
| 操作 | 耗时 | 带宽 | 备注 |
|---|---|---|---|
上传 put | 15.3s | ~66 MB/s | 写路径,含协议开销 |
下载 get | 4.9s | ~209 MB/s | 下载路径(GET),比上传快 3 倍 |
列表 ls | 0.25s | — | 元数据操作,毫秒级 |
💡 下载比上传快 3 倍:GET 是客户端视角的读路径——上传要走 RGW→OSD 三副本落盘,下载挑一份副本就够
- ⚠️ 但
s3cmd get含本地写盘开销(写到 /tmp),不是纯存储读 time三行里只看real(墙钟),user+sys是 CPU 耗时,不用管- 输出里的
[1 of 1]表示单文件直传、没走 multipart 分片;大文件(默认 >15MB 阈值)会自动拆成[part 1 of N]分片上传
文章分享
如果这篇文章对你有帮助,欢迎分享给更多人!















