You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
description: Use the version marker pattern to safely read and migrate stored data when a contract upgrade changes a data structure
5
5
---
6
6
7
-
When a contract is upgraded and a stored data structure gains new fields, the data already written to the ledger still uses the old layout. Naively reading those old entries with the new type causes the host to trap. This guide explains why that happens, introduces the version marker pattern as the correct solution, and covers lazy versus eager migration strategies and how to test them.
7
+
When a contract is upgraded and a stored data structure gains new fields, the data already written to the ledger still uses the old layout. Naively reading those old entries with the new type causes the host to trap. This guide introduces the version marker pattern as the correct solution, covers lazy versus eager migration strategies and how to test them, and explains why the "intuitive" approach fails.
8
8
9
-
## Why intuitive approaches fail
9
+
## Versioned Enum Pattern
10
10
11
11
Suppose a contract stores `DataV1` entries and is upgraded to use `DataV2`, which adds an optional field `c`:
12
12
13
13
```rust
14
14
#[contracttype]
15
-
pubstructData { a:i64, b:i64 }
15
+
pubstructDataV1 { a:i64, b:i64 }
16
16
17
17
#[contracttype]
18
18
pubstructDataV2 { a:i64, b:i64, c:Option<i64> }
19
19
```
20
20
21
-
### Approach 1: Read old entries directly with the new type
21
+
The recommended approach in this circumstance is to implement a versioned enum that can hold either a `V1` or `V2` data struct.
22
+
23
+
```rust
24
+
#[contracttype]
25
+
pubenumData {
26
+
V1(DataV1),
27
+
V2(DataV2),
28
+
}
29
+
30
+
#[contracttype]
31
+
pubenumDataKey {
32
+
Data(u64),
33
+
}
34
+
```
22
35
23
-
The most natural approach is to read the stored bytes directly as `DataV2` and expect `c` to default to `None`:
36
+
### Migration Logic
37
+
38
+
The migration logic enumerates the two data formats and converts `V1` data to `V2` format, and passes `V2` format through. If it's already `V1`, it maps fields `a` and `b` over and sets the new `c` field to `None` (the field that was added in `V2`). If it's already `V2`, it passes through unchanged. This is a lazy migration - old data is upgraded on read, not in a bulk migration.
24
39
25
40
```rust
26
-
// Reading a DataV1 entry with the DataV2 type.
27
-
// A developer might expect c = None for old entries — but this traps.
This traps with `Error(Object, UnexpectedSize)`. The Soroban host validates the field count of the XDR-encoded value against the type definition before returning anything to the contract. Because `DataV1` has two fields and `DataV2` has three, the host rejects the entry before the SDK can handle it.
51
+
### Reading with version awareness
33
52
34
-
### Approach 2: Use `try_from_val` as a fallback
53
+
The value is read from storage and then `into_v2()` ensures that the returned value is in the `V2` format.
35
54
36
-
Another approach is to use `try_from_val` expecting to catch a deserialization error and recover:
This also traps at the host level. The field count validation happens in the host environment during deserialization — it does not produce a Rust `Err` that the SDK can intercept. There is no way to catch or recover from the mismatch at the contract level.
72
+
### Testing migrations
50
73
51
-
The root issue is that a contract cannot determine which type an existing storage entry was written as just by reading it. That information must be stored explicitly.
74
+
Testing data migration requires simulating state written by an old contract version and verifying that the new contract reads it correctly.
75
+
76
+
In this test data in the `V1` format is first stored. Then it's read using the `read_data` function, which converts data in the `V1` format to V2 format with `into_v2()` before returning the result. The result is tested with `assert_eq!()`, and stored with the same `id` as it was stored with, which means the `V1` formatted data is overwritten with the same data in `V2` format.
77
+
78
+
Then the data is read from storage to verify it's stored in the `V2` format, and finally the data is read using the `read_data()` function to verify that the data is also returned in the `V2` format by the read function.
Data::V1(_) =>panic!("expected Data::V2 after write_data, found Data::V1"),
117
+
}
118
+
119
+
// Subsequent reads go through the V2 branch and return identical values.
120
+
letresult=client.read_data(&id).unwrap();
121
+
assert_eq!(result.a, 5);
122
+
assert_eq!(result.b, 6);
123
+
assert_eq!(result.c, None);
124
+
}
125
+
```
52
126
53
127
## Version Marker Pattern
54
128
55
-
The solution is to store a version number alongside each data entry, keyed by the same identifier. The contract reads the version first, then branches on the result to decode the payload with the correct type.
129
+
An alternative solution is to store a version number alongside each data entry, keyed by the same identifier. The contract reads the version first, then branches on the result to decode the payload with the correct type.
56
130
57
131
### Key layout
58
132
59
-
Define two variants in your key enum — one for the version marker and one for the payload — both keyed by the same `id`:
133
+
Define two variants in your key enum - one for the version marker and one for the payload - both keyed by the same `id`:
60
134
61
135
```rust
62
136
#[contracttype]
63
137
pubenumDataKey {
64
-
DataVersion(u32), // version marker keyed by id
65
-
Data(u32), // data keyed by id
138
+
DataVersion(u32), // version marker, keyed by id
139
+
Data(u32), // data, keyed by the same id
66
140
}
67
141
```
68
142
69
143
Each logical record occupies two storage slots. Because the version is stored per-record rather than globally, each entry is independently versioned. There is no all-or-nothing upgrade requirement.
70
144
71
145
### Reading with version awareness
72
146
73
-
Before decoding a storage entry, read its version marker. Use `unwrap_or(1)` to handle entries that were written before versioning was introduced — the absence of a version key is itself a signal that the entry is version 1:
147
+
Before decoding a storage entry, read its version marker. Use `unwrap_or(1)` to handle entries that were written before versioning was introduced. The absence of a version key is itself a signal that the entry is version 1:
74
148
75
149
```rust
76
150
fnread_data(env:&Env, id:u32) ->DataV2 {
@@ -105,7 +179,7 @@ Once version-aware read/write logic is in place, there are two strategies for co
105
179
106
180
#### Lazy migration (convert on read)
107
181
108
-
In lazy migration, old entries are left untouched on the ledger. When a record is read, its version is detected and it is up-converted in memory. When that record is later written back, it is stamped with the new version. No explicit migration step is needed — conversion happens as records are accessed in normal contract use.
182
+
In lazy migration, old entries are left untouched on the ledger. When a record is read, its version is detected and it is up-converted in memory. When that record is later written back, it is stamped with the new version. No explicit migration step is needed - conversion happens as records are accessed in normal contract use.
109
183
110
184
Lazy migration is generally preferred on blockchains. Leaving old entries untouched has no upfront cost and no risk of hitting instruction or ledger-entry limits at upgrade time. Records that are never accessed again are never migrated, which is usually acceptable.
Eager migration is rarely practical for large datasets on Soroban. Each rewrite consumes fees and burns instructions, and a single transaction cannot migrate an unbounded number of records — the contract will hit instruction or ledger-entry limits. If the batch must span multiple transactions, the contract is in a mixed-version state throughout the window, which means version-aware read logic is still required anyway.
210
+
Eager migration is rarely practical for large datasets on Soroban. Each rewrite consumes fees and burns instructions, and a single transaction cannot migrate an unbounded number of records - the contract will hit instruction or ledger-entry limits. If the batch must span multiple transactions, the contract is in a mixed-version state throughout the window, which means version-aware read logic is still required anyway.
137
211
138
212
Eager migration is occasionally appropriate when the total number of records is small and known in advance (for example, a fixed registry of a few dozen entries), or when you need to permanently drop old version branches from the read path.
// Read it — lazy migration produces a DataV2 in memory.
282
+
// Read it - lazy migration produces a DataV2 in memory.
209
283
letmigrated=read_data(&env, id);
210
284
assert_eq!(migrated.c, None);
211
285
212
-
// Write it back — this stamps the entry as version 2.
286
+
// Write it back - this stamps the entry as version 2.
213
287
write_data(&env, id, &migrated);
214
288
215
289
env.as_contract(&contract_id, || {
@@ -233,110 +307,45 @@ The three test cases cover the three states a record can be in after an upgrade:
233
307
- A `DataV2` entry written by the new contract
234
308
- A `DataV1` entry that is read and then written back (the lazy migration round-trip)
235
309
236
-
## Versioned Enum Pattern
237
-
238
-
Another approach is to implement a versioned enum that can hold either a `V1` or `V2` data struct.
239
-
240
-
```rust
241
-
#[contracttype]
242
-
pubenumData {
243
-
V1(DataV1),
244
-
V2(DataV2),
245
-
}
310
+
## Why intuitive approaches fail
246
311
247
-
#[contracttype]
248
-
pubenumDataKey {
249
-
Data(u64),
250
-
}
251
-
```
312
+
The techniques presented here may not immediately seem necessary. The "apparent" obvious solutions may be to programmatically handle the discrepancies in data types, rather than modify any of the underlying data structures, or adjust how the storage entries are read or written.
252
313
253
-
### Migration Logic
314
+
:::warning
254
315
255
-
The migration logic enumerates the two data formats and converts `V1` data to `V2` format, and passes `V2` format through. If it's already `V1`, it maps fields `a` and `b` over and sets the new `c` field to `None` (the field that was added in `V2`). If it's already `V2`, it passes through unchanged. This is a lazy migration — old data is upgraded on read, not in a bulk migration.
316
+
We've outlined a couple of the more "obvious" approaches to this problem, to illustrate _why_ these anti-patterns are not ideal. Please do not use the following code snippets as examples to be emulated. Rather, read the context of them, and learn why to avoid them.
256
317
257
-
```rust
258
-
implData {
259
-
pubfninto_v2(self) ->DataV2 {
260
-
matchself {
261
-
Data::V1(v1) =>DataV2 { a:v1.a, b:v1.b, c:None },
262
-
Data::V2(v2) =>v2,
263
-
}
264
-
}
265
-
}
266
-
```
318
+
:::
267
319
268
-
### Reading with version awareness
320
+
### Approach 1: Read old entries directly with the new type
269
321
270
-
The value is read from storage and then `into_v2()` ensures that the returned value is in the `V2` format.
322
+
You may think the most natural approach is to read the stored bytes directly as `DataV2` and expect `c` to default to `None`:
This traps with `Error(Object, UnexpectedSize)`. The Soroban host validates the field count of the XDR-encoded value against the type definition before returning anything to the contract. Because `DataV1` has two fields and `DataV2` has three, the host rejects the entry before the SDK can handle it.
280
333
281
-
The write function `write_data()` takes a data argument in the `DataV2` format.
334
+
### Approach 2: Use `try_from_val` as a fallback
335
+
336
+
Another approach is to use `try_from_val` expecting to catch a deserialization error and recover:
// This branch is never reached - the host traps before returning Err.
344
+
letv1=DataV1::try_from_val(&env, &raw).unwrap();
345
+
DataV2 { a:v1.a, b:v1.b, c:None }
286
346
}
287
347
```
288
348
289
-
### Testing migrations
349
+
This also traps at the host level. The field count validation happens in the host environment during deserialization - it does not produce a Rust `Err` that the SDK can intercept. There is no way to catch or recover from the mismatch at the contract level.
290
350
291
-
Testing data migration requires simulating state written by an old contract version and verifying that the new contract reads it correctly.
292
-
293
-
In this test data in the `V1` format is first stored. Then it's read using the `read_data` function, which converts data in the `V1` format to V2 format with `into_v2()` before returning the result. The result is tested with `assert_eq!()`, and stored with the same `id` as it was stored with, which means the `V1` formatted data is overwritten with the same data in `V2` format.
294
-
295
-
Then the data is read from storage to verify it's stored in the `V2` format, and finally the data is read using the `read_data()` function to verify that the data is also returned in the `V2` format by the read function.
Data::V1(_) =>panic!("expected Data::V2 after write_data, found Data::V1"),
334
-
}
335
-
336
-
// Subsequent reads go through the V2 branch and return identical values.
337
-
letresult=client.read_data(&id).unwrap();
338
-
assert_eq!(result.a, 5);
339
-
assert_eq!(result.b, 6);
340
-
assert_eq!(result.c, None);
341
-
}
342
-
```
351
+
The root issue is that a contract cannot determine which type an existing storage entry was written as just by reading it. That information must be stored explicitly.
0 commit comments