Footguns with Postgres “at time zone 'UTC'”
bookofrevenue.com156 points by birdculture a day ago
156 points by birdculture a day ago
This summer, with the help of AI, I found an inconsistency in the way Postgres handles timestamp vs. timestamptz comparisons under a DST spring-forward gap for the datetime_ops btree family [0]. Essentially there are scenarios where expression B > A and B < C, but also C = A, which can cause queries using a btree index (among other things) to return an incorrect result.
The assessment in the mailing list was that this was a bug, but there were no good ways to fix it.
> Backpatching a behavioral change like this seems awfully scary. For the moment I'm just contemplating what we could potentially change in master. So far I don't like any of the choices :-(
[0] https://www.postgresql.org/message-id/flat/CA%2BCOZaDmCuOds-...
The problem really is inherent to DST itself, just as the month math in TFA is inherently wonky in any system. What's January 30th + 1 month? February 28th (or 29th, if a leap year)? March 1st? March 2nd?
UI time elements have to be presented in the user's TZ. In the DB one should store timestamptz in UTC for all things, and maybe also timestamptz in non-UTC TZs for user input (e.g., in a calendaring app).
Does the same hole exist for range types? If a GiST exclusion constraint on tstzrange is built on the same comparison logic, a booking table could accept two reservations that overlap only inside the DST gap, and nobody would notice until two people show up for the same room. Curious whether you checked that path or only the btree family.
The text says: "Comparing a timestamp and timestamptz will always result in false."
That is wrong. Postgres changes the timestamp to timestamptz very quietly in the background. Wether it uses the time zone of the session for this. If it is true or false depends on the TimeZone setting. This is more bad than "always false". In production with UTC it works. On a laptop of a California developer it does not work.
Just tested:
SET TIME ZONE 'America/Los_Angeles';
SELECT '2026-03-01 00:00:00'::timestamp = '2026-03-01 00:00:00+00'::timestamptz; /* f */
SET TIME ZONE 'UTC';
SELECT '2026-03-01 00:00:00'::timestamp = '2026-03-01 00:00:00+00'::timestamptz; /* t */
If you can, use Postgres 16.
The double AT TIME ZONE 'UTC' thing from the text is not needed there. There you have date_add with a time zone as the third argument:date_add(b.month_start, interval '1 month', 'UTC')
This adds the month in UTC and it stays a timestamptz.
Automatic conversion using "global" state between "local"/"human" time (5pm where I am now) and points in time (ie with timezone) is one of the biggest sins imho for many libraries/db's/languages. Ran into it a bunch of times using C# as well.
The two most difficult things in software engineering...
Naming.
Timestamps.
You must not have invalidated your cache since a third thing.
Obviously there are more but the original joke circa 1999 was this. Made Y2K even more exciting.
> but the original joke circa 1999 was this.
No, the original joke, which GP is referring to, is “There are only two hard things in computer science. Naming things and cache invalidation.”
It predates 1999 and the Y2K bug by a fair amount. I first saw it on Usenet in the early 90s, around 1994 I think.
> There are only two hard things in computer science. Naming things and cache invalidation.
More like
There are only two hard things in
computer science. Naming, cache
invalidation, and off-by-one bugs.And of by one errors
You are off by one characters.
In my freshman Pascal class, I quickly learned how to use “iff” [if and only if] in a sentence, and my T.A. loved that.
It made sense when databases and programs were used almost exclusively locally. It still makes sense for local apps (e.g. local-first or local-only smartphone and desktop apps) who typically will automatically do the right thing that way based on the OS regional settings.
It only started causing widespread issues with the rise of cross-region internet SaaS. Database systems, language runtimes, and OS APIs are keeping the default behavior for backwards compatibility.
Yes and no, it was thought to make sense for "end-user-programmers" where it's helpful to be fully locale specific, I'm Swedish and my OS settings makes programs expecting comma (,) signs for decimal separation is something that's actually hit me today when copy-pasting between programs.
So in practice, while it was kinda useful to be locale/region dependant for some users it's probably been more trouble in the long run to be overly helpful.
C# was released after y2k, so they don't have the excuse.
Also, You're missing the biggest sin here however, locale specific time is OK, automatically allowing conversions/comparisons to points in time types without specifying timezones has in principle never caused anything but grief.
Author here
Wow, thank you for testing it. I did test it on my own laptop (Seattle).
> date_add(b.month_start, interval '1 month', 'UTC')
I didn't know this. Thank you.
The SQL standard is unfortunately really horrible when it comes to handling of time. The type `timestamp` is not a timestamp at all because it doesn't encode a unique point in time, it just stores a date and a time which has to be interpreted relative to a timezone. It should be called "datetime".
Moving a Java Instant back and forth between a database is also a surprisingly difficult task to do right, and it doesn't help that JDBC is just handling it completely wrong if you use its setTimestamp/getTimestamp methods. Not because it is a bad design with footguns, but because the implementation is just plain wrong and will corrupt your data if you deal with instants whose calendar date is far enough in the past due to it using the legacy date/time API which switches to the Gregorian calendar for dates in the past.
The name `timestamp with time zone` is also misleading because it doesn't actually store a time zone, it stores the number of seconds since epoch like a java.time.Instant (although at a different resolution). The "with time zone" part just refers to the textual format you denote the values in which includes the time zone after the date/time part to uniquely identify a timestamp, but the time zone is thrown away and not stored after the value has been parsed. This is different from e.g. `ZonedDateTime` in Java which will actually store the offset and therefore corresponds to a pair of (Instant, TimeZone).
A ZonedDateTime is not an instant and a timezone, for the same reason that you can’t unambiguously round trip between arbitrary timezones and UTC:
- zoned datetimes carry ambiguities as to their actual location on the timeline (because they can repeat, or not exist at all)
- future zoned datetime carry outright uncertainty as to their actual location on the timeline (as the zone’s offset can be updated at any point and any number of times until the event has elapsed)
Still, the GP is correct about the problems.
Relational databases have exactly 1 type that corresponds to modern data-handling practices: timestamp with time zone, that stores a timestamp. There is no good way to store any other modern type, and the 1980s practices on time handling weren't actually very good.
I'd say relational databases (in the sense of standard SQL) have 0 types that correspond to modern data-handling practices: timestamp with timezone stores an instant but lies about it and implicitly gets converted from and to the connection-local timezone.
It's always worth noting that future UTC timestamps are also ambiguous for certain operations, most notably computing durations, due to the unpredictability of leap seconds.
Has anyone proposed versioning timezones? Or is this such an edge case it would be overkill? (Either specify your future instant in UTC if you mean to stick to that, or specify it in a timezone and accept that it could change before it happens, or if you need something else get it in a contract and don't trust the computer!)
There is actually a simple heuristic you can use: if a point in time should be sticky to a calendar (e.g. calendar app or appointments which need to be synchronized between multiple humans or parties for a given context/location/region), store a datetime _without_ a timezone and make the timezone configurable for the user/infer it from the user. If you want a point in time which will not "physically" change, store a datetime _with_ a timezone, always, preferably UTC (e.g. logging, timers, measuring the occurrence of events).
The trouble is that you can't really make safe assumptions about whether to use the "sticky" paradigm or the "point-in-time" paradigm.
Sticky really only makes sense in two scenarios:
1. when all participants are assumed to be in the same geographic/political time zone for the foreseeable future (in which case the only advantage over point-in-time timestamps is future political changes to that region's time zone, like DST changes), or
2. when there's some privileged participant such that everyone else can assume events follow that participant's time zone (e.g. a company headquarters that moves very rarely, or an individual's personal wakeup alarms which can probably be assumed to follow their current location's time zone as they travel).
If you have a group of friends who like to stay in touch with regular group calls, and all/most of them are digital nomads who change their time zone of residence multiple times per year, you probably don't want the sticky paradigm.
In practice, if you have participants from different timezones, you agree on one "reference" timezone, which can be also UTC, as in your nomads example. You can consider UTC as just another calendar you can stick to (but you should just not hardcode this in software for this appointment use-case). You actually need a reference to be able to plan an event in the first place. Otherwise you don't know which time you can propose. You would need to propose a "fixed" time for every timezone which participates, but that doesn't work, because these times might not refer to the same physical instant, so the meeting would not be (fully) synchronized. So instead, you take some time from some timezone and translate that to all the other timezones. You can even see that in online games, where some of them have an official "server time", which is globally the same, helping players to meet at the same time instant. UTC itself is another examples of this, used e.g. for global navigation or air traffic control.
That "only advantage" is doing a lot to dismiss the actual use case of, say, everyone keeping appointments on their calendar when the legislature passes laws around time zone changes.
What’s the difference between these two other than how the client would convert to display to a user?
The 2 different types of datetimes encode different information and is not purely a display/presentation issue, it's also storage issue.
I happen to call them "scientific datetime" vs "cultural/political datetime". However, the software dev industry has not converged on a standard vocabulary to delineate the 2 types which is unfortunate because that means programmers are unaware that the difference exists. Concepts are more top-of-mind when there are good names to label them.
If a programmer doesn't understand how the 2 datetimes behave differently, they will create software bugs as I've outlined before: https://news.ycombinator.com/item?id=39418897
There's the meme of "store UTC everywhere" (maybe perceived as correct because of superficial similarity to "use UTF-8 everywhere") ... but storing datetimes as UTC is only unambiguous for historical events such as timestamps of activity stored in server logs.
But future datetimes can have ambiguous edge cases which causes the split into 2 different types.
What you call "cultural/political datetime" should be "standard time" (or maybe "civil time"):
https://en.wikipedia.org/wiki/Standard_time
But this is still different to a time someone enters into a calendar. Standard time can change its offset (to UTC) over time (e.g. DST), while a time in a calendar is fixed in the nominal sense.
"scientific datetime" is quite ambiguous, since I would consider science-level precision time to be TAI (International Atomic Time, what UTC uses as a reference), or maybe UT1, which is one variant of UT (Universal Time, unrelated to UTC), depending on the scientific field. For simple cases, UTC might be enough, so you could call this "UTC".
I think practically what matters for developers are three things:
- Standard time (dependent on timezone)
- UTC (the reference for standard times in the different timezones)
- Calendar times (seems to be called "floating time" [0]), just referring to a specific date and time, usually independent from both standard time and UTC, from the author's perspective (others viewing a foreign calendar might see times interpreted in their own timezone). Often scoped by physical location, but not necessarily.
, while a time in a calendar is fixed in the nominal sense.
The above scenario of fixed time regardless of DST/TZ changes is what I tried to call "cultural/political time". In other comments, I called it "appointment time".
What you call "calendar time", others will call it "time with calculated UTC offset". (Which then leads to more meta discussion of "no... calendar time is not UTC offset because ..." )
Both examples of our ambiguous labels causing more confusion is prime example of the industry not converging on good names to make devs aware of the difference.
>I think practically what matters for developers are three things:
That categorization is fine but is still obscuring the key issue: many developers think they can collapse all of your 3 types into one simple strategy of "always store it as UTC"
It's not only about display. If you store a user's appointment only as a UTC timestamp, you actually can't know the hour of the day (and the day itself to be precise) on which this appointment should happen, for a given calendar (probably the user's calendar, in a specific non-UTC timezone). You would have to guess by using the calendar's/user's timezone and compute some offset with UTC. But what if the user changes timezones or the timezone itself changes its value? Store without a timezone, and you know the exact hour and day the user intended. One is pointing to a day and hour in a calendar, the other is pointing at a point on the line of a linear timeline.
You're still just talking about client conversion on the write instead of the read. A Unix epoch is a UTC timestamp. That number is not time-zone-less it's just in the default computer time zone.
I'm referring to a "datetime", which should be a data type which stores a date, e.g. "2026-09-28", and a time, e.g. "16:58:18". The unix epoch is not relevant here, unless you _want_ to store a UTC timestamp. You could store a UTC timestamp also as ISO 8601, which is not a number. But this is not related to the timezone-less datetime I'm talking about. To make my points more clear, just imagine the datetime is stored as a string "2026-09-28 16:58:18" with no implied timezone whatsoever.
But there is zero use case difference between these two things you mentioned:
> if a point in time should be sticky to a calendar store a datetime _without_ a timezone
> If you want a point in time which will not "physically" change, store a datetime _with_ a timezone
Both of these things are exactly identical. You are still storing an exact point in time in both cases. The only thing about a calendar use case is the presentation layer.