问题简介
restore point卡住15分钟报错waiting for schema info finishes reloading
环境说明
v8.5.4
三节点混布
问题详细说明
需要部署ticdc,使用br restore point进行数据初始化,还原失败报错,即使重试也是失败,错误相同。
只进行restore full可以成功。
[2026/09/10 10:03:57.079 +08:00] [ERROR] [main.go:42] ["br failed"] [error="failed to wait until schema reload: waitUntil timed out after waiting for 15m0s"] [errorVerbose="waitUntil timed out after waiting for 15m0s\ngithub.com/pingcap/tidb/br/pkg/utils.WaitUntil\n\t/workspace/source/tidb/br/pkg/utils/wait.go:42\ngithub.com/pingcap/tidb/br/pkg/task.waitUntilSchemaReload\n\t/workspace/source/tidb/br/pkg/task/stream.go:1904\ngithub.com/pingcap/tidb/br/pkg/task.restoreStream\n\t/workspace/source/tidb/br/pkg/task/stream.go:1502\ngithub.com/pingcap/tidb/br/pkg/task.RunStreamRestore\n\t/workspace/source/tidb/br/pkg/task/stream.go:1243\ngithub.com/pingcap/tidb/br/pkg/task.RunRestore\n\t/workspace/source/tidb/br/pkg/task/restore.go:701\nmain.runRestoreCommand\n\t/workspace/source/tidb/br/cmd/br/restore.go:75\nmain.newStreamRestoreCommand.func1\n\t/workspace/source/tidb/br/cmd/br/restore.go:249\ngithub.com/spf13/cobra.(*Command).execute\n\t/root/go/pkg/mod/github.com/spf13/cobra@v1.8.1/command.go:985\ngithub.com/spf13/cobra.(*Command).ExecuteC\n\t/root/go/pkg/mod/github.com/spf13/cobra@v1.8.1/command.go:1117\ngithub.com/spf13/cobra.(*Command).Execute\n\t/root/go/pkg/mod/github.com/spf13/cobra@v1.8.1/command.go:1041\nmain.main\n\t/workspace/source/tidb/br/cmd/br/main.go:40\nruntime.main\n\t/usr/local/go/src/runtime/proc.go:272\nruntime.goexit\n\t/usr/local/go/src/runtime/asm_amd64.s:1700\nfailed to wait until schema reload"] [stack="main.main\n\t/workspace/source/tidb/br/cmd/br/main.go:42\nruntime.main\n\t/usr/local/go/src/runtime/proc.go:272"]
解决方案
结论
https://github.com/pingcap/tidb/pull/66092
Problem Summary: Without calling session.SetSchemaLease, BR will use the default schema lease duration of 1s which is only intended for testing, not production. When the PD TSO somehow lags the BR wall clock by over 0.5s (1/2 of schemaLease) it will not be able to leave the schema reload loop.
由于集群之间时钟不统一(虽然配置了chronyc 时钟同步,但是远端时钟Server不通,导致差距在1s以上),触发br的bug。
修复方案
1)调整时钟,确保集群间各个节点时钟一致。
2)将br暂时升级为v8.5.8版本,即可正常基于 restore point 的还原。
查看br 版本 : tiup br --version
br的升级方式:如果可联网直接tiup install br:v8.5.8 即可;离线环境可以只升级br组件,单独下载指定版本的br,替换二进制文件或调用时直接指定新版本br的位置。
失败和成功日志对比
对比成功和失败日志,正常情况下加载schema信息不到1ms。
[tidb@tidb-01 tempdir]$ grep reloading restore_point_20260910_*log
-- br 8.5.4
restore_point_20260910_fail.log:[2026/09/10 09:48:56.916 +08:00] [INFO] [stream.go:1899] ["waiting for schema info finishes reloading"]
-- br 8.5.8
restore_point_20260910_succ.log:[2026/09/10 11:54:59.221 +08:00] [INFO] [stream.go:2225] ["waiting for schema info finishes reloading"]
restore_point_20260910_succ.log:[2026/09/10 11:54:59.221 +08:00] [INFO] [stream.go:2233] ["reloading schema finished"] [timeTaken=1.347µs]
到达该步骤,下游数据其实已经写完了,失败的只是后续步骤(CleanUpKVFiles / InsertGCRows / RepairIngestIndex 等步骤)。
拓展
chronyc tracking 逐项判读
字段 |
当前值 |
判读 |
Reference ID |
761F2863 () |
上游 NTP 源 |
Stratum |
3 |
正常 |
Ref time (UTC) |
Tue Jul 28 07:27:13 2026 |
最后一次有效测量距今约 44 天 |
System time |
0.000001631 s fast |
这是「最后一次测量 + 频率补偿」的外推值,非当前实测 |
Last offset |
-0.002225280 s |
44 天前那次测量的偏移 |
RMS offset |
0.001312303 s |
历史统计,非当前 |
Frequency |
9.560 ppm fast |
内核已补偿的晶振频偏 |
Residual freq |
-0.342 ppm |
补偿后的残余频偏,会持续累积 |
Skew |
1.818 ppm |
频率不确定度(上界) |
Root dispersion |
11.971472740 s |
累积离散度约 12 秒(健康集群通常 < 1s) |
Update interval |
1042.3 s ≈ 17 分钟 |
轮询间隔,与陈旧 Ref time 直接矛盾 |
重点看是否有 ^* 行;全 '?' 或 'x' = 未同步
shell> chronyc sources -v
210 Number of sources = 2
.-- Source mode '^' = server, '=' = peer, '#' = local clock.
/ .- Source state '*' = current synced, '+' = combined , '-' = not combined,
| / '?' = unreachable, 'x' = time may be in error, '~' = time too variable.
|| .- xxxx [ yyyy ] +/- zzzz
|| Reachability register (octal) -. | xxxx = adjusted offset,
|| Log2(Polling interval) --. | | yyyy = measured offset,
|| \ | | zzzz = estimated error.
|| | | \
MS Name/IP address Stratum Poll Reach LastRx Last sample
===============================================================================
^? 113.141.164.38 3 10 0 45d -16ms[ -17ms] +/- 153ms
^? 118.31.3.89 2 10 0 45d +264us[-2096us] +/- 41ms
如上^?代表着时钟不可达,说明了各个节点时钟差异的原因。