My Backup Strategy Finally Grew Up

For over ten years I had a personal backup strategy. Strictly speaking it didn’t cover every scenario I could encounter, but the plus side was I had one at all, and it saved me a few times despite the gaps.

The tool I used was Resilio Sync. Basically a BitTorrent style sync system for personal use, replicating data bidirectionally between wherever I told it to go. At various points that meant a copy on my laptop’s external hard drive, a copy on my NAS, more copies scattered around the house on Raspberry Pis, and for a while, a copy on a VPS. Before that I even had encrypted copies sitting at two different friends’ houses, encrypted specifically so they couldn’t see what was in them, but I could still recover if something went wrong.

It was great, for what it was. It solved exactly one use case: drive failure, or something happening to the house. And it worked. When my NAS died I was sure for a minute I’d lost everything, and then remembered a huge chunk of that data was safely sitting somewhere else entirely.

I paid for Resilio once, something like ten years ago, roughly fifty dollars, plus a twenty dollar upgrade at some point. That’s the entire lifetime cost I paid, and I got my money’s worth many times over. At one point I tried Syncthing, an open source alternative that does something similar, just to see if free was as good as paid. It was fine, but fiddly enough that I didn’t really stick with it back then. More on that later, because that part of the story isn’t over.

The hardware I used for this setup evolved constantly underneath it. I started on Raspberry Pi 2s next to the NAS, moved to Pi 3s, then Pi 4s. I’ve blogged about this system in various forms over the years, so if you go digging through the archive you’ll find the whole lineage.

But here’s the thing I knew for almost as long as I ran it: it never covered data corruption. Not really. Early on that was a vague worry about my own data getting mangled somehow. As cyber threats got worse, and they always were getting worse, it turned into a much sharper question. Was I actually prepared to survive ransomware, or a wiper attack wiping out everything the replication system would happily sync to every node?

So on top of the replication, I’d take manual offline snapshots. Once a year if I was disciplined, once a quarter if I was lucky. Sometimes that meant another spot on the NAS, which isn’t really a proper backup since it’s still one box, but at least it was off the live replication system. Other times it meant a physical hard drive, especially for directories too big to replicate offsite, like my media library.

Generally though, I wasn’t that worried. Everyone’s a target for that kind of attack in theory, but my actual threat model used to be drive failure, full stop. That’s what kept me up at night, if anything did. But the ransomware / wiper concern kept creeping up over time, the way it does.

About a year or two ago I finally decided to address the other backup use cases and to take into account having an immutable backup. I set up my own Borg backup on a Raspberry Pi. I did it mostly manually, with some AI help for suggestions along the way. Then starting around March or April this year, I got a lot more serious about building things fully AI assisted, with the AI actually managing systems directly over SSH. That shift naturally pointed straight at my backup setup. I kept interrogating the AI on the best way to build it out, and as I was standing up new websites and other projects, the whole thing evolved with them.

A couple of iterations later, here’s where it landed. Two separate Borg backup targets. One local, in the house, backing up the big media directories and everything else internally. Good for an immutable copy, but obviously useless if something happens to the house itself, which is why I still take a manual physical hard drive backup periodically. They’re sitting in a safe right now. Future Scott needs to actually get them somewhere outside the house.

Feeding into that, I also moved my laptop’s sync off Resilio and onto the NAS, and funnily enough ended up back on Syncthing to do it. Not because it got better, but because the AI found it easier to set up and manage than Resilio. I set up one way replication, so anything that changes on the NAS doesn’t sync back to the laptop, and that gets picked up on the NAS side for the real backup. It worked perfectly on the first try, which after years of fighting with sync tools felt almost suspicious.

The second target is a VPS, two terabytes, that I got as a deal from lowendbox.com for about thirty dollars a year. Ridiculous deal. Every VPS and Raspberry Pi I run backs up daily to a directory on the NAS. The NAS then spins up a Docker instance for each Borg destination, runs the backup, and shuts itself back down. There’s a separate prune job on my laptop I run to clean up old backups over time.

It’s honestly more robust than a personal backup needs to be. I’ve run a fair number of threat models on it at this point, probably more than necessary, but that’s the improvement over the old system: it’s immutable now, not just replicated.

I’ve got a pattern set up with the AI now where I can just ask it to check the status of everything and give me a detailed report. One of the VPSs also connects into my home network over Tailscale to monitor status, with alerting wired up, so I get a push notification the moment anything isn’t working right.

Really robust. More than I need. Definitely worthwhile anyway.

The current iteration has been running for two months at time of writing this in late August 2026. There was an earlier version before it that ran about two months, before I rebuilt it with some lessons learned. So far, so good.

I’m still debating whether to ship a Raspberry Pi and a large external drive to a friend’s house to replicate everything at home, offsite, the old school way. Then I do the math and realize the total cost of that probably outweighs several years of the two terabyte VPS backup I’m already paying for. So probably not worth it. Still evaluating though, because apparently doing ROI calculations on my own backup strategy is a hobby now, service versus hardware, one spreadsheet at a time.

Overall, very pleased with where it landed.

Gigabit FOMO

After my recent internet upgrades I ended up with a new dilemma. A good dilemma, but still.

For a while now I have been running Tailscale on all my devices. It handles all my internal services and lets me connect to everything on my home network from anywhere without having to think about it. For privacy on top of that I added the Mullvad VPN exit nodes that Tailscale offers as an add-on. Five dollars a month, integrated straight into the Tailscale network, pick a city and go. No toggling anything off, no conflicts with my internal services. It has been a great setup.

Then I upgraded my home internet to gigabit. And suddenly I started noticing something. On my phone, without the VPN, I was pulling 800 to 900 megabit down. With the VPN on, more like 500. And something in my brain did not like that.

I kept telling myself it was probably in my head. But I wanted to actually know for sure.

So I did some research and thought it through properly. And the conclusion was pretty straightforward. On a phone, do I actually need gigabit download speeds? Am I going to notice the difference between those numbers when loading a webpage? The answer is no. At those speeds the bottleneck is never my connection. It is the server on the other end, or the screen, or just how fast things can render. The gigabit matters when I am downloading something substantial on my computer or my NAS. On a phone or tablet, even streaming HD video, it simply does not matter.

What I had was not a performance problem. It was a psychological one. I upgraded to gigabit, I could see a faster number without the VPN, and some part of my brain decided I was being shortchanged even though in practice I was not missing anything.

The rational answer is obvious. Keep the VPN on. It is more private, more secure, and the speed difference is not something I will ever notice in real use. I just needed to actually verify that before I could accept it.

VPN stays on.

AI For the Mundane Stuff…

I have been using AI for increasingly mundane things I always meant to do but never actually got around to. Most recently, a mystery Mac DMG file.

I had a Mac DMG file. About eight gigs. Password protected. I have no idea why I password protected it. I knew roughly what was in it, a big archive of installer packages going back years, but I could not get in because the password I thought it was was not working.

I have been saving installers since back when software did not just update itself over the internet and you actually needed to hold onto major versions. Over time the archive grew across different Mac architectures. PowerPC stuff, then x86 and Universal binaries, and now arm64. So there is a lot in there going back a long time.

I put Claude on the case. I gave it my best guesses at what the password might be, because I did not want it going through every possible combination until the end of time, and it set up a cracking tool locally on my MacBook Air and let it run with those hints. It got in.

So now I had all this data, plus a bunch of other old archives, and the problem was figuring out what was what. File dates only get you so far. Some of it was obvious, certain versions of Microsoft Office are a dead giveaway for the era and architecture. But a lot of it was not.

This matters because I have two PowerPC Macs I picked up vintage that I actually want to use, and I have been thinking about picking up an older Intel Mac for similar reasons. Half of it is just because it is cool to have. The other half is that occasionally I do want to run some old software. So getting the right installer on the right machine is not a trivial question.

A few prompts later, Claude had written a set of scripts that parsed through everything, sorted it by architecture, and organised it all into the right folders. A few back and forth sessions where it asked me questions and I confirmed things, and it had a pretty solid result.

Then I went a step further. I had it size and create read only DMG images for each architecture, correctly sized so the PowerPC stuff would fit on a DVD if I want to burn one, or on a USB key I can move to one of the older machines. Right file systems, right sizes, the lot. About 100 gigs of stuff in total so it took a while to churn through, but my actual involvement was maybe 20 to 30 minutes of back and forth.

I had wanted to do this for years. Total elapsed time was maybe three hours, almost all of it the computer doing the work.

Now it is all backed up, the images are ready, and whenever I get around to spinning up those old machines I can actually do it. Which might not be for a while. But at least now I can.

My Watch Is All I Need (For Payments)

When metal credit cards first came out, I thought they were cool. They were novel. They looked premium. They were basically indestructible compared to a regular plastic card. And at the time, only a handful of cards even had them, so there was something a bit different about having one.

I was a little vain about it. Fine.

But I have the opposite opinion now. And the reason is pretty simple. I do not really carry a wallet anymore. I have one, but it lives in my bag as emergency disaster recovery. We are talking a card with a decent amount of credit for emergencies if something happens to the phone or watch, and a card or two for cash withdrawals.

In the UK I almost never touch a physical card anyway. My phone and my watch handle everything. Apple Pay on both. Tap and go. For years now the only reason I pull out a card at all is to get cash from an ATM, and even that is rare.

We did an entire trip to France last year without once pulling out a physical card or touching cash. The whole thing. Did not miss it.

So when a subscription I pay for decides the premium perk they are going to give me is a metal card, I am not impressed. I am just thinking about how two metal cards in that slim wallet means I can’t actually use it as a backup the one time I actually need it.

Physical card use is down. Cash use is down. The trajectory is pretty obvious.

I have left the house with just my watch and been completely fine. Done it more than once. It works.

Going fully watch-only as a daily habit is probably a stretch. But the wallet is already basically an afterthought. Metal cards are so 2018. Current and future Scott is all in on the watch and phone combo.

Finally Wired

I’ve wanted to wire our house properly since we moved in and were renting it. When we bought it I figured I’d finally get around to it. That was a while ago.

The first real progress came in 2022 when we had work done and the electrician ran two CAT 6 cables into the floor before they poured the cement. Great start. Except he didn’t actually connect them to anything useful. The most logical endpoint was under the stairs, where the main power is and where you can get cables upstairs or to the back of the house pretty easily. The internet comes into the house at the media centre in the living room. He didn’t run cable there. I kind of wanted him to. I don’t really know why I didn’t push harder on that. After the main scope of work was done, it was hard to get a hold of him anyway, so I just forgot about it and dropped the subject.

So those cables sat there, coiled in the wall, for a few years.

Fast forward to recently. We were finally ripping out the carpeting on the first floor. It had been there since before we moved in eight years ago and it just needed to go. While the floors were up, I had them trunk a CAT 6 run down to under the stairs. Progress.

Still needed someone to do the last run to the living room, which is the most important bit. I had two different electricians doing other work on the house at the time. Both said yeah, sure, we can do that. Neither followed through. I ended up on Checkatrade, found someone slightly overpriced, didn’t care, hired them. They went out the side of the house and back in rather than under the floor, which isn’t exactly what I wanted, but it’s honestly how the old CAT 3 phone cable was done back in the day and it works fine. They came in right next to where the Openreach FTTP fibre enters the house, which was actually convenient.

That was early May. Finally.

Once the cables were in, I plugged in the two TP-Link Deco units I had. I’ve always used my own router rather than letting the Decos handle routing. I just don’t like all-in-one devices. The interfaces never let you do what you actually want. The Decos were two or three generations old at that point but they were decent. Testing the wired connection I was getting close to a gigabit on the cables. On wireless I was seeing around 600Mbps, and on the wired connection upstairs, over 900. Not bad.

Then I added a third Deco in the family room at the back of the house, which has always been a dead zone. My wife would sit by the window back there and we basically couldn’t use the internet. Moved one of the older units there, put the newer one upstairs in the office. Immediately better.

That’s when I started thinking about the internet plan.

The wireless backhaul on the old Decos was using the 6GHz band, which meant the 5GHz and 2.4GHz bands were doing all the work for actual devices. It was a bottleneck. I was only getting 330Mbps download from a 350Mbps plan because so much capacity was eaten up by the backhaul. Now that the access points were wired in, that bottleneck was gone. So why was I still on a 350Mbps plan?

I could upgrade to 550Mbps, or just go straight to gigabit. The gigabit plan is about £20 more a month than what I was paying. Triple the speed. Most of our devices can’t fully saturate a gigabit connection, but enough of them can that it seemed worth it. I upgraded the Decos to Wi-Fi 7 while I was at it. Not the latest generation. That was too expensive. But one back from that. I swapped out the old gigabit switch under the stairs for a 2.5G one. That means everything except the 16-port gigabit switch in the media centre is now 2.5G. I’m not too worried about the media centre switch. It’s mostly feeding Raspberry Pis that aren’t anywhere near saturating a gigabit uplink.

I put the order in and literally the next day my speed test app on Docker was showing 920Mbps down. Thank you A&A. I run a speed test every hour to keep an eye on things and I’m consistently getting 920 to 927Mbps down and around 106 to 108Mbps up. Compare that to the 295Mbps I was getting before.

The only weak spot is the loft. It has to get signal through the floor from the office below, and while I’m still getting around 300Mbps up there, running a cable all the way up the side of the house doesn’t seem worth it. We don’t spend much time up there anyway.

My neighbour and I were geeking out about the fact that I get 800Mbps in the middle of the back garden.

I knew. I just didn’t.

Photo is an old one from my “start-up” days. Partsearch one of the Layer 3 switch’s as we were putting it in the call center. I thought it was appropriate!

The ISP That Operates Like a Tech Person Designed It

When I was building out my OPNsense router, I had a decision to make. Keep running Pi-hole for DNS and DHCP, or switch over to the built-in AdGuard on the router itself.

I’d been using Pi-hole for a while and the functionality was genuinely compelling. But if I was critical of my own uses, I wasn’t really using most of it. The main thing I got out of the Pi-Hole was ad blocking at the DNS level, and I can get that from a privacy-focused DNS provider like Mullvad anyway. So I made the call to consolidate everything onto the router.

Turns out that decision had a nice side effect.

My ISP is Andrews & Arnold, and they give me an IPv6 block along with a dedicated IPv4 address for my router. There were some limitations with Pi-hole when it came to properly handing out public IPv6 addresses to devices on the network. Could be a Pi-hole thing. Could be me never having quite sorted it. I am not really sure. Either way, moving everything to the router cleared it up. Now every device on my network gets its own real, internet-facing IPv6 address.

Do I need this? Absolutely not. But the network engineer version of me from the early 2000s would have been beside himself.

Which brings me to Andrews & Arnold, because I don’t think I’ve written about them directly before and they are pretty awesome. Also it feels like they are a bit unique in terms of internet providers.

I first used them when we moved to the UK and I needed a DSL connection. They were great, but the Openreach copper line quality at the time was a problem and the speeds weren’t there, so I ended up switching to Virgin for a few years just to get something faster. Then Openreach rolled out FTTP to our neighbourhood and I went straight back to A&A.

They are not outrageous, but they are noticeably more expensive than the value options out there. When I first signed up for FTTP I had to pick a usage tier, which I have a moral objection to on principle, and I ended up needing to go above the one terabyte plan pretty quickly. To their credit they’ve since moved their higher tier to unlimited, so that’s sorted itself out eventually.

But here’s why I pay it. Their documentation literally has sections like “want to run your own DNS server? Go ahead, we don’t mind.” When you place an order you have to acknowledge that you’re getting completely unfiltered internet access. Your static IP address comes on a little plastic card in the box like a hotel key, which I know is just a piece of plastic, but the confidence of that gesture still lands.

I was at a talk a while back where someone was explaining the steps they’d gone through to get their own IPv6 block registered in their own name rather than their ISP’s. They they went on to discuss their ISP assigning it for their use at their house. The process sounded painful and involved. I could tell from the context it was a UK audience, and I quietly suspected I knew which ISP they were using. Afterwards I asked the guy. Yep, Andrews & Arnold.

I priced out EE recently for comparable or even higher bandwidth plans. Substantially cheaper. But online the answers about what restrictions they actually impose are, charitably, murky. With Andrews & Arnold I know exactly what I’m getting. And what I’m getting is an ISP that operates like a tech person designed it. From what I have read that is because it was founded by tech people.

Still well worth the money.

Collecting AI Models Like It’s a Hobby

It’s almost like I collect AI models and providers. Each one has its own use case, and I probably could consolidate. But in some situations it’s not even worth the effort to do it.

As of writing this in early May 2026, and I mention the date because I usually have a backlog of a couple of months before things go live, my main tool is Claude. Well worth the money and then some.

But in addition to that I’m paying for Venice.AI as my privacy AI. Very cheap per year, limited use case, but for what it does it’s worth every penny.

Then there’s Lumo from Proton. I’m about one week into an experiment of not using it to see if I actually miss it. The honest answer is probably not, because I’m now running local models on my new MacBook Air with 32 gigs of RAM, which I got specifically to be able to do this. I’ve been running Mistral 3 3B locally and early results suggest it might actually be more private and more capable than Lumo for what I need it for. So that £120 a year might be on its way out.

I also have Perplexity Pro, which I don’t actually pay for directly. It comes bundled with another subscription I have. It’s slightly better than the others for web search type queries so I use it for that. Would I pay for it separately? Probably not. But I have it so I use it.

And then there’s ChatGPT. I touched on this in the last post but I’m currently trialling whether running Claude Code and ChatGPT’s Codex side by side is actually cheaper than paying for the next Claude tier up. The overages on Claude Code have been real, and the maths might just work out. Still figuring that one out.

So yes, I have more models and services than I probably need. Lumo is the obvious cut. Beyond that I’m genuinely not sure. The ecosystem as a whole is delivering value right now, even if it’s a bit sprawling.

What I do know is that two years in, I’m using these tools in ways I couldn’t have imagined when I first paid for Copilot on a 30 day trial just to see what would happen. And the people I talk to who haven’t really dug in yet, I get it, it can feel overwhelming. But the gap between what you can do with these tools and what most people are doing with them is pretty wide right now. And that gap is only going to get wider.

Levelling Up to Claude Code

So I mentioned at the end of my last post about AI that switching to Claude unlocked a new level of what I was doing with it.

It started small. Someone at work mentioned they’d been using Claude Code to build actual apps, and they’re not a developer either. Just someone with ideas and enough curiosity to see what happens. That stuck with me.

I had a specific problem I wanted to solve. I have some VoIP numbers through my internet provider, actual UK mobile numbers I can receive texts on. Useful for giving out to LinkedIn contacts or anyone I don’t want having my main number. The problem was sending replies meant logging into a clunky website every time. Nobody wants to do that.

So I asked Claude about it. One thing led to another and it said it could help me build a web app for that. When I asked if we could run it in Docker it said of course. I already run Docker on a few systems at home so that made sense. It built me a container, I ran it on my Synology, got it working for one number, cloned it for a second. Then I thought, wait, can we just have one container with a dropdown to switch between numbers? Of course we can. And it did it.

That was the moment I thought, if it can build me an app with this level of input and time, what else can it really do?

So I set up Claude Code properly, gave it access to my Raspberry Pi setup, and had it take the SMS app further. It interrogated my existing config, made some enhancements, deployed everything from the command line. Straightforward, but impressive.

From there I got more ambitious. My blog had been running on YunoHost, which is a self-service VPS platform. Decent enough, but it’s always on an older version of Debian because the open source volunteers take what feels like a long time to update it, and if an app isn’t in their package store you’re out of luck. I’d always wanted to run my own properly configured stack but never wanted to deal with the time to care and feed it.

I had a spare VPS sitting around doing nothing. So I asked Claude, can we design and build my entire website on this thing. WordPress in Docker, proper backups, the works. It said yes.

First I designed the whole thing with Claude, got a proper design document together, then imported that into Claude Code and let it build. About $25 to $40 in API costs later I had a website. And not just a website. Automated daily backups following a daily, weekly, monthly cadence with pruning built in, all replicating to a second location with an immutable copy at the end. Backup infrastructure honestly better than some hosting providers I’ve heard about. I then migrated my site over to it, wiped the old setup, and migrated it back again just to prove I could restore it. All worked.

Then I got bold. I had the spare VPS now freed up, so I used it to build a set of personal tools I’d always wanted but never had the time to set up properly. An open source SSO system with passkey authentication for me and my wife. A Searx search instance sitting behind the SSO. Network monitoring. I didn’t even know I could set up my own SSO until Claude walked me through it. Some API costs and an afternoon later it was running.

Around the same time I set up a proper monitoring stack. My external VPS now watches my home internet connection and my main Raspberry Pi, and sends push notifications to my phone if anything goes down. Not email, because I’m not going to look at an email. An actual push notification. I also have a speed test running every hour on a gigabit ethernet connection straight into the router, and my internet provider, who I’ve never had a bad word to say about, is consistently delivering around 305 megabit on a 350 megabit plan. Seeing that graphed over time is genuinely satisfying.

All of that cost me some API charges, probably a couple of dollars worth, and £4 for a push notification app I now own outright. Not a subscription. Just mine.

Then I built a router.

I’d looked at this before. A router is basically just a computer with extra network cards, and I’d used open source router software in the past. But I’d always bought the manufacturer’s hardware because I didn’t want to deal with building and maintaining my own. This time I asked Claude to help me design and deploy it, and then help me maintain it going forward.

Of course It said yes. It can be a yes man/lady/person a lot.

I researched the hardware with Claude’s help and landed on a Protectli VP2430, a fanless little box with four 2.5 gigabit network ports and an Intel N150 processor. 16 gigs of RAM and a 256 gig SSD. Overkill for a router, which means it’ll last a long time. Then I designed the whole OpnSense configuration with Claude before touching any hardware.

Deployment was more painful than I expected. The biggest issue was needing the new router connected to my laptop while also needing live internet on the same laptop to use Claude to programme it. The address space decisions I’d made complicated things further. I spent six hours one evening and an entire Saturday on it before rolling back. Then I realised what was in the end the hurdle that had stopped me, made one design adjustment, tried again the following week, and had it running in about two hours. It’s been stable since.

I’m not a software developer. But I’ve spent years building and supporting data centres, call centres, and large scale applications, so I understand enough to know when something looks wrong and push back on it. I understand some scripting and basic fundamentals. What I can do is explain what I want clearly, spot when the output doesn’t smell right, and interrogate it until it does. That combination, it turns out, is enough to build some pretty serious stuff.

The things I’ve been able to do aren’t things I couldn’t have done before in theory. But the time to care and feed a self-hosted setup was always more than I was willing to put in. Now I have AI that helps me design it right, build it right, and fix it when something goes wrong. The barrier that used to stop me isn’t really there anymore.

So now I’m looking at what’s next. I have mockups for a couple of actual apps I want to build. Things I want for myself that don’t exist quite the way I want them. A colleague at work just builds whatever he thinks of now. I’m getting there.

The biggest current headache is cost. Claude Code API charges add up fast when you’re doing serious work, and last month I had enough overages that I’m now trialling whether running both Claude Code and ChatGPT’s Codex together is actually cheaper than paying for the next Claude tier up. Early signs are interesting.

And separately, with my new MacBook Air running 32 gigs of RAM, I’m finally in a position to run proper local models. I’ve started downloading and testing, and the early results suggest I might be able to replace some of what I’m using Lumo for with a local model that’s actually more private and possibly better. That’s an ongoing experiment.

It’s a lot. But honestly, talking to people now, whether friends, colleagues, or people in the industry, so many are either just scratching the surface or not looking at it at all. I was over a year late getting serious about this. I don’t think I’m late anymore.

Switching to Claude

I’ve been on this AI journey for a while now, and for most of it ChatGPT was my main tool. But in February 2026 that changed.

Someone whose opinion I trust recommended I give Claude a try after voicing my frustrations with ChatGPT. When I actually sat down to evaluate it I did something that in hindsight was pretty funny. I asked Claude directly why I should use Claude over ChatGPT.

It told me not to.

Based on what I told Claude I wanted, it said I’d probably get better results from ChatGPT. So naturally I didn’t trust that answer and kept pushing. As I interrogated it further it started explaining that it was slower and more thoughtful, and from its previous read of what I was looking for, it figured I wanted fast straight answers. And that’s when something clicked for me.

One of my biggest frustrations with ChatGPT was exactly that. I’d ask it something and it would just fire back an answer. Fast, confident, and often not what I actually asked for. Like if I asked for specific instructions on how to do something on an Apple product, it would give me generic steps that didn’t even exist in the actual interface. I’d have to stop it and say don’t give me fluff, give me the actual thing. Then it would have to go look it up and either admit it didn’t know or finally give me something useful.

What Claude was describing as a weakness, slower and more considered, was exactly what I wanted. So I signed up and started using both in parallel.

Early on I was genuinely impressed with what was coming out of Claude. I was using Sonnet, the middle tier model, and the difference in output quality was noticeable pretty quickly. The concern then was whether I’d end up paying for two services. I was already paying for Venice and Lumo on the privacy side and the last thing I wanted was more sprawl.

But it became clear fairly fast that Claude was where I wanted to be. Which meant I had to migrate everything I’d built in ChatGPT over the previous six months or so. Custom GPTs, saved prompts, all of it. I had to extract everything, make sure I had backups, build a little text based database of all my prompts, and systematically move it all across.

I got it done within a month and managed to avoid paying for both services at the same time for more than one month. Then I downgraded ChatGPT to the free tier and haven’t looked back.

From a reliability standpoint Claude is better. Not perfect, and everything I said in the last post about not trusting it still applies. But it’s a meaningful improvement. And honestly, switching to it unlocked a whole new level of what I started doing with AI. Which is what the next post is about.

Use AI Like It’s Lying To You, Because It Is

I’ve touched on this in passing across a few posts now, but it deserves its own space. Because as useful as AI has been for me, it is not sunshine and rainbows.

AI is a great tool. I genuinely believe that. But I also think the future gets pretty dystopian if people don’t use these tools with their eyes open. And right now, a lot of people aren’t.

One thing I read recently that stuck with me: it can do super advanced calculations but it can’t tell time right. That sounds like a joke but it isn’t. Some of the things I’ve asked it, it will be absolutely insistent it’s correct. You interrogate it because something smells off, and eventually it folds. Oh yeah, you’re right, I was wrong. I’ve had situations where I knew it was wrong, kept pushing, and it took a surprisingly long time before it admitted it.

So what I tell people, my kids, colleagues at work, is this. Use it. But do not trust it. Assume it’s going to lie to you. If you go in with that mindset and you scrutinise the output, it can be really good. But you have to be able to scrutinise it. That’s the part people skip.

That’s also the part that makes it genuinely dangerous in the wrong hands. I can ask it something about cybersecurity and I’ll know pretty quickly if the answer looks right or completely off. But I can’t ask it to do my taxes. I don’t know tax law. So if it tells me I can do something, I have no idea if it’s true. That’s a problem. And it’s why you see things like lawyers submitting court filings with citations that don’t exist because a judge caught them using AI and not checking the output. People just throw stuff in and take whatever comes out.

I ran into this myself about a year or so ago when I was doing some budget planning. Nothing super sensitive, just the savings pots I set aside for predictable expenses throughout the year so I’m not hit with an inconsistent spend later. I do it all in a spreadsheet and it gets involved. I figured let me see if AI can handle it.

It did it, and then it didn’t. The numbers were inconsistent. Flat out wrong in places. I tried for a few months and I just could not rely on it. So I stopped and went back to my spreadsheet.

More recently I tried again with a different model, and I’ll get into that in the next post. But it’s actually working now. Two months in and the output is consistently accurate. I’ve also gotten smarter about how I prompt it, asking it to show its work and export the data in a way I can verify. So it’s a combination of the models improving and me getting better at using them.

But the overall lesson hasn’t changed. The reading of the tea leaves is the hard part. Sometimes the output is exactly what you wanted and better than you could have done yourself. Sometimes it’s close but slightly off in a way that’s easy to miss. And sometimes it’s just wrong and completely confident about it.

The tool is getting better. That’s real. But so is the risk of people treating it like it’s infallible. It isn’t. Not even close.