0
0
0
0
博客/.../

br restore point报错waiting for schema info finishes reloading

 克里克里克  发表于  2026-10-08

问题简介

restore point卡住15分钟报错waiting for schema info finishes reloading

环境说明

v8.5.4

三节点混布

问题详细说明

需要部署ticdc,使用br restore point进行数据初始化,还原失败报错,即使重试也是失败,错误相同。

只进行restore full可以成功。

[2026/09/10 10:03:57.079 +08:00] [ERROR] [main.go:42] ["br failed"] [error="failed to wait until schema reload: waitUntil timed out after waiting for 15m0s"] [errorVerbose="waitUntil timed out after waiting for 15m0s\ngithub.com/pingcap/tidb/br/pkg/utils.WaitUntil\n\t/workspace/source/tidb/br/pkg/utils/wait.go:42\ngithub.com/pingcap/tidb/br/pkg/task.waitUntilSchemaReload\n\t/workspace/source/tidb/br/pkg/task/stream.go:1904\ngithub.com/pingcap/tidb/br/pkg/task.restoreStream\n\t/workspace/source/tidb/br/pkg/task/stream.go:1502\ngithub.com/pingcap/tidb/br/pkg/task.RunStreamRestore\n\t/workspace/source/tidb/br/pkg/task/stream.go:1243\ngithub.com/pingcap/tidb/br/pkg/task.RunRestore\n\t/workspace/source/tidb/br/pkg/task/restore.go:701\nmain.runRestoreCommand\n\t/workspace/source/tidb/br/cmd/br/restore.go:75\nmain.newStreamRestoreCommand.func1\n\t/workspace/source/tidb/br/cmd/br/restore.go:249\ngithub.com/spf13/cobra.(*Command).execute\n\t/root/go/pkg/mod/github.com/spf13/cobra@v1.8.1/command.go:985\ngithub.com/spf13/cobra.(*Command).ExecuteC\n\t/root/go/pkg/mod/github.com/spf13/cobra@v1.8.1/command.go:1117\ngithub.com/spf13/cobra.(*Command).Execute\n\t/root/go/pkg/mod/github.com/spf13/cobra@v1.8.1/command.go:1041\nmain.main\n\t/workspace/source/tidb/br/cmd/br/main.go:40\nruntime.main\n\t/usr/local/go/src/runtime/proc.go:272\nruntime.goexit\n\t/usr/local/go/src/runtime/asm_amd64.s:1700\nfailed to wait until schema reload"] [stack="main.main\n\t/workspace/source/tidb/br/cmd/br/main.go:42\nruntime.main\n\t/usr/local/go/src/runtime/proc.go:272"]

解决方案

结论

https://github.com/pingcap/tidb/pull/66092

Problem Summary: Without calling session.SetSchemaLease, BR will use the default schema lease duration of 1s which is only intended for testing, not production. When the PD TSO somehow lags the BR wall clock by over 0.5s (1/2 of schemaLease) it will not be able to leave the schema reload loop.

由于集群之间时钟不统一(虽然配置了chronyc 时钟同步,但是远端时钟Server不通,导致差距在1s以上),触发br的bug。

修复方案

1)调整时钟,确保集群间各个节点时钟一致。

2)将br暂时升级为v8.5.8版本,即可正常基于 restore point 的还原。

查看br 版本 : tiup br --version

br的升级方式:如果可联网直接tiup install br:v8.5.8 即可;离线环境可以只升级br组件,单独下载指定版本的br,替换二进制文件或调用时直接指定新版本br的位置。

失败和成功日志对比

对比成功和失败日志,正常情况下加载schema信息不到1ms。

[tidb@tidb-01 tempdir]$ grep reloading restore_point_20260910_*log

-- br 8.5.4

restore_point_20260910_fail.log:[2026/09/10 09:48:56.916 +08:00] [INFO] [stream.go:1899] ["waiting for schema info finishes reloading"]

-- br 8.5.8

restore_point_20260910_succ.log:[2026/09/10 11:54:59.221 +08:00] [INFO] [stream.go:2225] ["waiting for schema info finishes reloading"]

restore_point_20260910_succ.log:[2026/09/10 11:54:59.221 +08:00] [INFO] [stream.go:2233] ["reloading schema finished"] [timeTaken=1.347µs]

到达该步骤,下游数据其实已经写完了,失败的只是后续步骤(CleanUpKVFiles / InsertGCRows / RepairIngestIndex 等步骤)。

拓展

chronyc tracking 逐项判读

字段

当前值

判读

Reference ID

761F2863 ()

上游 NTP 源

Stratum

3

正常

Ref time (UTC)

Tue Jul 28 07:27:13 2026

最后一次有效测量距今约 44 天

System time

0.000001631 s fast

这是「最后一次测量 + 频率补偿」的外推值,非当前实测

Last offset

-0.002225280 s

44 天前那次测量的偏移

RMS offset

0.001312303 s

历史统计,非当前

Frequency

9.560 ppm fast

内核已补偿的晶振频偏

Residual freq

-0.342 ppm

补偿后的残余频偏,会持续累积

Skew

1.818 ppm

频率不确定度(上界)

Root dispersion

11.971472740 s

累积离散度约 12 秒(健康集群通常 < 1s)

Update interval

1042.3 s ≈ 17 分钟

轮询间隔,与陈旧 Ref time 直接矛盾

重点看是否有 ^* 行;全 '?' 或 'x' = 未同步

shell> chronyc sources -v

210 Number of sources = 2

.-- Source mode '^' = server, '=' = peer, '#' = local clock.

/ .- Source state '*' = current synced, '+' = combined , '-' = not combined,

| / '?' = unreachable, 'x' = time may be in error, '~' = time too variable.

|| .- xxxx [ yyyy ] +/- zzzz

|| Reachability register (octal) -. | xxxx = adjusted offset,

|| Log2(Polling interval) --. | | yyyy = measured offset,

|| \ | | zzzz = estimated error.

|| | | \

MS Name/IP address Stratum Poll Reach LastRx Last sample

===============================================================================

^? 113.141.164.38 3 10 0 45d -16ms[ -17ms] +/- 153ms

^? 118.31.3.89 2 10 0 45d +264us[-2096us] +/- 41ms

如上^?代表着时钟不可达,说明了各个节点时钟差异的原因。

0
0
0
0

版权声明:本文为 TiDB 社区用户原创文章,遵循 CC BY-NC-SA 4.0 版权协议,转载请附上原文出处链接和本声明。

评论
暂无评论