Some advice from someone who's done it wrong for years and dealt with dead people who have done it wrong for years.
Firstly don't try and run a corporate level backup solution or NAS with terabytes of storage and tape at home. You're going to drop dead one day and everyone will hate you for what you leave them to deal with. They will have no idea how to pick through the remains to get important stuff out. It took me a whole 3 days, as someone who knows how this works, to get into a dead relative's NAS appliance and find any mention of any info to get into his EV which had basically bricked itself in the driveway while he was in hospital.
Secondly, delete everything. Not joking. Delete as much as you can. Everything you've watched, delete it. All the old emails, delete them. Tracks you don't like on CDs you've ripped, delete them. Old software ISOs, delete them. Thousands of photos of the same thing taken 1 second apart, delete them. Thousands of PDFs of books you will never read, delete them. I went down from a whopping 8TB to 300Gb doing that and it's still shrinking monthly. Eventually it'll fit on one mediocre computer and that makes life so so so much easier.
Then look at your backup strategy. It'll look simple then. Mine is three external disks.
1. 1TB T7 shield (not encrypted) which lives at my partner's house and gets an rsync once a month.
2. 1TB T7 shield (encrypted) which lives in my bag and gets an rdiff-backup once a week.
3. 2TB Lacie HDD (not encrypted) which lives in my fire safe at home and gets an rdiff-backup once a week.
If I drop dead, my partner can just plug the thing into her computer and get stuff off it.
Don't make your life any more complicated than it needs to be. I implore you.
> The size of Sync is actually fairly small for me, it's just 12GB, but it's not small enough to fit on my 128GB phone.
Synctrain is a great iOS syncthing client that connects to your swarm but doesn't actually sync any files (unless you ask it to). You can still browse the contents of your shared folders and selectively download individual files.
One thing I would add to a modern backup strategy: a deferred offline copy
Sure it's bothersome to use the sneakernet once a month (manually copy data to USB stick and stash it somewhere), but when that ransomware hits (can happen on personal devices also), it literally can't access the offline copy, and try to encrypt that one also.
Immutable backup solutions are code and thus eventually defeatable from the remote attackers point of view. Offline requires physical access, and maybe a big enough wrench to get the USB stick's disk encryption PIN out of you.
Tape + Iron Mountain is difficult to beat for offline copies.
Surrendering just a little bit of absolute control over the physical media can dramatically increase the likelihood that it survives whatever apocalyptic event. I don't know that I've ever heard of tapes being stolen from Iron Mountain. Even if this happened, what are the odds that they would get all the tapes. Your tapes? It probably looks like the warehouse from Raiders of the Lost Ark in there.
I have a locker at my place of work where I store a few HDDs and USBs. These are the most up to date, but I also have others at my parents place and my in-laws. Gets troublesome keeping track of which ones are up to date as of what date. Good challenge for staying organised though, I've got a whole naming system and numbered hierarchy and scripts that run ordered by priority.
I use a similar system. I have paper console tape on each drive and write the date last used on it and also keep a text file log of which drive and when. The backup batch file also writes a timestamp.txt to the root of the drive.
I also have a set I curate and put in with the family & household papers. That doesn't have any of my personal crap, just documents, photos, and family history etc.
I think there is a system to do this with git (maybe git ostree or annex) that will just keep the hash link in a repo but have the file in a offline drive, so you can easily keep track of files that are in cold storage. It may need you to have your own encryption system though.
I do this too, but with ZFS. So, since the snapshots have the creation timestamp in their name, it's obvious which drive has the latest data. But I also tend to remember if I went to the office or to my parents' house last.
Immutable backup solutions are code and thus eventually defeatable from the remote attackers point of view.
Yeah, but an attacker is unlikely to compromise both you and your storage provider at the same point. Many S3-compatible storage providers (e.g. Backblaze and Hetzner) support object lock, where you can lock objects for a certain number of days (and refresh locks if the objects are still used in recent backups). Typically object locks are implemented such that not even you yourself can remove the data from the account settings.
E.g. when I cancelled my Backblaze B2 account, I had to wait until the object locks expired before I could remove the remaining data and delete the account.
Arq on macOS has great support for object locks BTW.
Ideally an offline copy that can be made read-only with a physical switch. So that when you're trying to restore, you know nothing is going to mess with it. Not sure what the right solution is for something like that.
> One of the other materials that I could not figure out how to back up properly is emails. The reason is that there is no clear way to back it up systematically
Doesn't that depend entirely on the storage format? For example if you use thunderbird as a client it has database files that you can copy. Since you say that one of your devices is running arch you should check out some of the email options in the repos. There are at least a few of options that can act as a middleman to bridge between your external accounts and your clients.
> So having a plan to roll back to a stable state might be helfpul. I installed and set up Timeshift for this
Snapshots make this trivial. Particularly in the case of btrfs if you configure your bootloader to always use the default subvolume then rollback becomes just `btrfs subvolme set-default`.
imap-backup. It just extracts everything from imap server as mbox. Then that's rdiff-backup'ed to target disks. You can imap-backup it back to another IMAP server if you need to.
However it's better to just delete most stuff so you don't have to back it up. Think my mailbox is about 50Mb.
1. Important things I care about: these go into a folder that gets backed up to Hetzner (examples include photos in Immich, code pushed to Forgejo, etc)
2. Smaller files I sort-of care about: these usually just go into Proton Drive (examples include design documents)
3. Large files I don't care about losing: these go into a folder on a hard drive that is not backed up (examples include media)
This took me about a day to set up and I just check in on it every once in a while. Once a year I will do a practice disaster recovery run.
What does your practice recovery run look like? I've got backup set up, but I don't know how to best test it. For example, do you restore everything from Hetzner, or some random sample?
I don't like these staged backup for these reasons. Every step (e.g. NAS backup to Hetzner) can cause an error, so adding steps increases the probability of error. Personally, I just back up directly from my machines to a local storage server and to a cloud object storage with object lock (so that e.g. ransomware encryption or attempts to remove the data do not work).
What the author describes here is hard because it's "simple."
What's simple because it's "hard" is replacing parts 2 & 3 with a network appliance like TrueNAS running a zfs pool that syncs to backblaze every night. Yeah you have to learn a bit but it won't fall in weird ways like the hard drive part here will just fail to mount one night and not back things up for 3 months until you notice. My 2¢
> like the hard drive part here will just fail to mount one night and not back things up for 3 months until you notice
I think the author is having trouble because he is conflating concepts and roles that should be distinct. Sync, rotating snapshots, and deduplicated backups need to be kept entirely separate if you want any hope of maintaining your sanity.
So he's got sync but he's missing some sort of rotating snapshot system which would solve the stated concern of guarding against syncthing replicating corrupted data. Such automated snapshots can then be used as the source to feed the backup pipeline.
That hard drive doesn't make a good backup because it seems that it is always online. You need an offline backup that you manually plug in to run the job once every so often.
He's also making this more difficult than it needs to be by insisting that the backup drive be compatible with windows. Plug the drives for both snapshots and backups into a linux box, format them with a modern filesystem, and get on with life.
Sync is its own clusterfuck and I have yet to arrive at a satisfactory solution myself despite wasting inordinate amounts of time on it. IMO you either go with a network share or you make due with the "least bad" option of syncthing. Personally I've more or less settled on sshfs at this point not because it's particularly good but because it works well enough and doesn't add any additional complexity.
Personally I use btrfs snapshots on all my devices, those get streamed across the network to a NAS, and the contents of the NAS are periodically (every few months) stuffed into borgbackup on redundant offline devices. Aside from sync the other problem you'll run into if you're a data hoarder is how to split backups across multiple drives once you exceed a few TB. Because external drives only get so large but the NAS will inevitably keep ballooning.
> Scaling to hundreds of thousands of files is not a problem, scaling beyond that and git will start to get slow.
So that's probably insufficient for me by at least a couple orders of magnitude. I'm able to maintain my sanity because snapshots simply capture device state, the NAS collects all snapshots while maintaining their independence, and (so far) borg has been sufficiently scalable to deduplicate any collection I've thrown at it.
I agree I think one of the main things I learned from all this was that I should probably buy / set up a real NAS. I’ll look into zfs pool thanks for the comment!
This caused all sorts of grief at day job where dedupe was useful for a big cache. Conversely I've run ZFS at home for like 15 years at this point without trouble. But this one is an absolute nightmare.
It sounds like most of your frustration can be eliminated by actually using your NAS as the ground-truth for all data, using something like SMB/NFS instead of SyncThing.
Oh interesting I hadn’t considered this. I guess the only trouble would be if I am out with my laptop and I have no internet connection but this seems like a good tradeoff for simplicity
Beware of interactions with git and syncthing. It's fine for straightforward repos where all you do is commit but as soon as you start doing more complicated branching and re-basing you will start to generate lots of `sync-conflict` files. I haven't really found a reliable way around this so I've decided to just manually rsync from my desktop onto my laptop when I want to work remotely (or more recently, SSH into my desktop directly instead and work off that).
Apple's Time Machine still backs up your Mac even when you don't have access to the backup target (USB isn't connected, network isn't available). It just stores the backup information on your local storage and then transfers it over to the actual backup target when it's back online. Don't know if there's similar solutions for non-macOS systems though.
> It contains at the home directory the folder Sync which is what gets synced across all devices and what needs to be maintained.
That's a core Dropbox-level mistake of putting the backup cart before the main use horse - your files should be sorted according to their primary, not backup use
> Syncthing works very well, its only constraint is that it does need the devices to be often on if you want consistent syncing.
So it's working very poorly, this is a very inconvenient constraint
> phone is a whole other beast. It does not have enough storage to
Indeed
Overall, that fits the title perfectly, and is a bad way to setup your backups. Special needs need special apps.
Some advice from someone who's done it wrong for years and dealt with dead people who have done it wrong for years.
Firstly don't try and run a corporate level backup solution or NAS with terabytes of storage and tape at home. You're going to drop dead one day and everyone will hate you for what you leave them to deal with. They will have no idea how to pick through the remains to get important stuff out. It took me a whole 3 days, as someone who knows how this works, to get into a dead relative's NAS appliance and find any mention of any info to get into his EV which had basically bricked itself in the driveway while he was in hospital.
Secondly, delete everything. Not joking. Delete as much as you can. Everything you've watched, delete it. All the old emails, delete them. Tracks you don't like on CDs you've ripped, delete them. Old software ISOs, delete them. Thousands of photos of the same thing taken 1 second apart, delete them. Thousands of PDFs of books you will never read, delete them. I went down from a whopping 8TB to 300Gb doing that and it's still shrinking monthly. Eventually it'll fit on one mediocre computer and that makes life so so so much easier.
Then look at your backup strategy. It'll look simple then. Mine is three external disks.
1. 1TB T7 shield (not encrypted) which lives at my partner's house and gets an rsync once a month.
2. 1TB T7 shield (encrypted) which lives in my bag and gets an rdiff-backup once a week.
3. 2TB Lacie HDD (not encrypted) which lives in my fire safe at home and gets an rdiff-backup once a week.
If I drop dead, my partner can just plug the thing into her computer and get stuff off it.
Don't make your life any more complicated than it needs to be. I implore you.
> The size of Sync is actually fairly small for me, it's just 12GB, but it's not small enough to fit on my 128GB phone.
Synctrain is a great iOS syncthing client that connects to your swarm but doesn't actually sync any files (unless you ask it to). You can still browse the contents of your shared folders and selectively download individual files.
https://apps.apple.com/us/app/synctrain/id6553985316
One thing I would add to a modern backup strategy: a deferred offline copy
Sure it's bothersome to use the sneakernet once a month (manually copy data to USB stick and stash it somewhere), but when that ransomware hits (can happen on personal devices also), it literally can't access the offline copy, and try to encrypt that one also.
Immutable backup solutions are code and thus eventually defeatable from the remote attackers point of view. Offline requires physical access, and maybe a big enough wrench to get the USB stick's disk encryption PIN out of you.
Tape + Iron Mountain is difficult to beat for offline copies.
Surrendering just a little bit of absolute control over the physical media can dramatically increase the likelihood that it survives whatever apocalyptic event. I don't know that I've ever heard of tapes being stolen from Iron Mountain. Even if this happened, what are the odds that they would get all the tapes. Your tapes? It probably looks like the warehouse from Raiders of the Lost Ark in there.
I have a locker at my place of work where I store a few HDDs and USBs. These are the most up to date, but I also have others at my parents place and my in-laws. Gets troublesome keeping track of which ones are up to date as of what date. Good challenge for staying organised though, I've got a whole naming system and numbered hierarchy and scripts that run ordered by priority.
Well overdue for a refresh.
I use a similar system. I have paper console tape on each drive and write the date last used on it and also keep a text file log of which drive and when. The backup batch file also writes a timestamp.txt to the root of the drive.
I also have a set I curate and put in with the family & household papers. That doesn't have any of my personal crap, just documents, photos, and family history etc.
I think there is a system to do this with git (maybe git ostree or annex) that will just keep the hash link in a repo but have the file in a offline drive, so you can easily keep track of files that are in cold storage. It may need you to have your own encryption system though.
I do this too, but with ZFS. So, since the snapshots have the creation timestamp in their name, it's obvious which drive has the latest data. But I also tend to remember if I went to the office or to my parents' house last.
Immutable backup solutions are code and thus eventually defeatable from the remote attackers point of view.
Yeah, but an attacker is unlikely to compromise both you and your storage provider at the same point. Many S3-compatible storage providers (e.g. Backblaze and Hetzner) support object lock, where you can lock objects for a certain number of days (and refresh locks if the objects are still used in recent backups). Typically object locks are implemented such that not even you yourself can remove the data from the account settings.
E.g. when I cancelled my Backblaze B2 account, I had to wait until the object locks expired before I could remove the remaining data and delete the account.
Arq on macOS has great support for object locks BTW.
Ideally an offline copy that can be made read-only with a physical switch. So that when you're trying to restore, you know nothing is going to mess with it. Not sure what the right solution is for something like that.
> One of the other materials that I could not figure out how to back up properly is emails. The reason is that there is no clear way to back it up systematically
Doesn't that depend entirely on the storage format? For example if you use thunderbird as a client it has database files that you can copy. Since you say that one of your devices is running arch you should check out some of the email options in the repos. There are at least a few of options that can act as a middleman to bridge between your external accounts and your clients.
> So having a plan to roll back to a stable state might be helfpul. I installed and set up Timeshift for this
Snapshots make this trivial. Particularly in the case of btrfs if you configure your bootloader to always use the default subvolume then rollback becomes just `btrfs subvolme set-default`.
imap-backup. It just extracts everything from imap server as mbox. Then that's rdiff-backup'ed to target disks. You can imap-backup it back to another IMAP server if you need to.
However it's better to just delete most stuff so you don't have to back it up. Think my mailbox is about 50Mb.
or use imapsync to pull emails, then backup the resulting folder.
I have a server in a closet with a few hard drives attached to it. All of my devices connect to it via Tailscale.
I use the wonderful https://github.com/garethgeorge/backrest as a web UI around restic.
Every night, I back the data up to a Hetzner storage box https://www.hetzner.com/storage/storage-box/ which is only ~$3.50/mo USD for 1TB of data.
I have three "tiers" of data for myself:
1. Important things I care about: these go into a folder that gets backed up to Hetzner (examples include photos in Immich, code pushed to Forgejo, etc)
2. Smaller files I sort-of care about: these usually just go into Proton Drive (examples include design documents)
3. Large files I don't care about losing: these go into a folder on a hard drive that is not backed up (examples include media)
This took me about a day to set up and I just check in on it every once in a while. Once a year I will do a practice disaster recovery run.
What does your practice recovery run look like? I've got backup set up, but I don't know how to best test it. For example, do you restore everything from Hetzner, or some random sample?
I don't like these staged backup for these reasons. Every step (e.g. NAS backup to Hetzner) can cause an error, so adding steps increases the probability of error. Personally, I just back up directly from my machines to a local storage server and to a cloud object storage with object lock (so that e.g. ransomware encryption or attempts to remove the data do not work).
What the author describes here is hard because it's "simple."
What's simple because it's "hard" is replacing parts 2 & 3 with a network appliance like TrueNAS running a zfs pool that syncs to backblaze every night. Yeah you have to learn a bit but it won't fall in weird ways like the hard drive part here will just fail to mount one night and not back things up for 3 months until you notice. My 2¢
> like the hard drive part here will just fail to mount one night and not back things up for 3 months until you notice
I think the author is having trouble because he is conflating concepts and roles that should be distinct. Sync, rotating snapshots, and deduplicated backups need to be kept entirely separate if you want any hope of maintaining your sanity.
So he's got sync but he's missing some sort of rotating snapshot system which would solve the stated concern of guarding against syncthing replicating corrupted data. Such automated snapshots can then be used as the source to feed the backup pipeline.
That hard drive doesn't make a good backup because it seems that it is always online. You need an offline backup that you manually plug in to run the job once every so often.
He's also making this more difficult than it needs to be by insisting that the backup drive be compatible with windows. Plug the drives for both snapshots and backups into a linux box, format them with a modern filesystem, and get on with life.
Sync is its own clusterfuck and I have yet to arrive at a satisfactory solution myself despite wasting inordinate amounts of time on it. IMO you either go with a network share or you make due with the "least bad" option of syncthing. Personally I've more or less settled on sshfs at this point not because it's particularly good but because it works well enough and doesn't add any additional complexity.
Personally I use btrfs snapshots on all my devices, those get streamed across the network to a NAS, and the contents of the NAS are periodically (every few months) stuffed into borgbackup on redundant offline devices. Aside from sync the other problem you'll run into if you're a data hoarder is how to split backups across multiple drives once you exceed a few TB. Because external drives only get so large but the NAS will inevitably keep ballooning.
> Sync, rotating snapshots, and deduplicated backups need to be kept entirely separate if you want any hope of maintaining your sanity.
Not if you use git-annex.
Huh, I'd heard the name before but I hadn't realized how capable it was. Unfortunately when it comes to a data hoarder such as myself:
https://git-annex.branchable.com/scalability/
> Scaling to hundreds of thousands of files is not a problem, scaling beyond that and git will start to get slow.
So that's probably insufficient for me by at least a couple orders of magnitude. I'm able to maintain my sanity because snapshots simply capture device state, the NAS collects all snapshots while maintaining their independence, and (so far) borg has been sufficiently scalable to deduplicate any collection I've thrown at it.
I agree I think one of the main things I learned from all this was that I should probably buy / set up a real NAS. I’ll look into zfs pool thanks for the comment!
I liked the post because it tells a true story about one of the remaining problems that are hard to solve well without a 3rd party.
Fair warning with ZFS: do not turn on deduplication at the moment.
There is a straight up data loss bug in 2.4.3 (zeroed out files, totally silent).
https://github.com/openzfs/zfs/issues/18366
This caused all sorts of grief at day job where dedupe was useful for a big cache. Conversely I've run ZFS at home for like 15 years at this point without trouble. But this one is an absolute nightmare.
It sounds like most of your frustration can be eliminated by actually using your NAS as the ground-truth for all data, using something like SMB/NFS instead of SyncThing.
Then you only need to backup the NAS.
Oh interesting I hadn’t considered this. I guess the only trouble would be if I am out with my laptop and I have no internet connection but this seems like a good tradeoff for simplicity
Beware of interactions with git and syncthing. It's fine for straightforward repos where all you do is commit but as soon as you start doing more complicated branching and re-basing you will start to generate lots of `sync-conflict` files. I haven't really found a reliable way around this so I've decided to just manually rsync from my desktop onto my laptop when I want to work remotely (or more recently, SSH into my desktop directly instead and work off that).
I have a ~/git folder which I keep bare git repos in and push to.
That gets synced and it's been trouble free so far.
Apple's Time Machine still backs up your Mac even when you don't have access to the backup target (USB isn't connected, network isn't available). It just stores the backup information on your local storage and then transfers it over to the actual backup target when it's back online. Don't know if there's similar solutions for non-macOS systems though.
https://support.apple.com/en-us/102154
Exactly.
Just get a tape drive. Then you will find out how nice and simple disks are
> It contains at the home directory the folder Sync which is what gets synced across all devices and what needs to be maintained.
That's a core Dropbox-level mistake of putting the backup cart before the main use horse - your files should be sorted according to their primary, not backup use
> Syncthing works very well, its only constraint is that it does need the devices to be often on if you want consistent syncing.
So it's working very poorly, this is a very inconvenient constraint
> phone is a whole other beast. It does not have enough storage to
Indeed
Overall, that fits the title perfectly, and is a bad way to setup your backups. Special needs need special apps.
What solution do you use that syncs/backs up while the device is turned off? IME stuff?
Why wouldn't you use NTFS for backing up a windows system?
It's a pain in the ass on new systems if it copied restrictive permissions to reset/gain access.
Ntfs equivalent of Chmod/chown is a "go have a long lunch" type of operation.