I was curious as to where some Threadiverse instances were hosted, as not all explicitly indicate this. Thought that I'd also mention this here in case someone else has an interest in same.
I'll use lemmy.world as an example, as it's the most-popular Threadiverse instance and doesn't presently declare a location (at least via whatever mechanism is propagated to lemmy.fediverse.observer; I assume that fediverse.observer uses the Lemmy API to query it, though I've not dug into the matter).
Historically, one could resolve a hostname to an IP address and then could get the approximate location of an IP address in a number of ways; using whois will give the location of the netblock owner of the IP of a instance, which is generally good enough if one just cares about the approximate region, to find an instance near one.
However, the increasing use of CloudFlare for various reasons, including caching and DDoS prevention, has meant that often, that just provides the closest CloudFlare server to the requestor:
$ host lemmy.world|head -n1
lemmy.world has address 104.26.9.209
$ whois 104.26.9.209|egrep "^(NetName|Country)"
NetName: CLOUDFLARENET
Country: US
$
I'm in the San Francisco Bay Area; that'll just return a San Francisco CloudFlare netblock IP.
It is possible to measure latency to a server, which is often good enough to get a general idea of where that server is, unless the server is severely overloaded. One can't use, say, ICMP to reach a server behind CloudFlare, but one can use HTTPS.
$ httping -c3 -lg https://lemmy.world/
PING lemmy.world:443 (/):
connected to 172.67.71.35:443 (1075 bytes), seq=0 time=452.89 ms
connected to 104.26.9.209:443 (1081 bytes), seq=1 time=403.00 ms
connected to 104.26.9.209:443 (1080 bytes), seq=2 time=464.63 ms
https://lemmy.world/ ping statistics
3 connects, 3 ok, 0.00% failed, time 4321ms
round-trip min/avg/max = 403.0/440.2/464.6 ms
The minimum time there is what we're looking for, since the server cannot respond before it gets our request. But...this also is setting up a TLS session for each request, which is a lot of overhead that's going to increase the latency. Better to use HTTP keepalives to cut session setup time out of the equation:
$ httping -c3 -Qlg https://lemmy.world/
PING lemmy.world:443 (/):
pinged host 104.26.9.209:443 (1097 bytes), seq=0 time=437.16 ms
pinged host 104.26.9.209:443 (1084 bytes), seq=1 time=329.65 ms
pinged host 104.26.9.209:443 (1087 bytes), seq=2 time=332.28 ms
https://lemmy.world/ ping statistics
3 connects, 3 ok, 0.00% failed, time 4135ms
round-trip min/avg/max = 329.7/366.4/437.2 ms
$
Okay, but...it also takes time to generate the Web page, which adds overhead. We'd like to use a webpage that is fast to generate to minimize that. Plus, it's impolite to add load to instances by hammering an instance requesting expensive-to-generate pages repeatedly. So we'd like to find a webpage that's pretty cheap to generate and doesn't get cached by CloudFlare.
For example, lemmy.world's server thumbnail is static and CloudFlare can cache it and serve it from a local CloudFlare server, so any latency we measure is just going to be latency to CloudFlare:
$ wget --server-response https://lemmy.world/pictrs/image/0fd47927-ca3a-4d2c-b2e4-a25353786671.png -O/dev/null 2>&1|grep cf-cache-status
cf-cache-status: HIT
$
I remember seeing some discussion about the /setup URL being exposed on Lemmy instances. It's non-functional once the server has been set up (or it'd be a security risk), but it's not cached by CloudFlare and while I haven't gone and tried to benchmark it, I expect that it's cheap for an instance to generate, as there's virtually no content there:
$ wget --server-response http://lemmy.world/setup -O/dev/null 2>&1|grep cf-cache-status
cf-cache-status: DYNAMIC
$
So using that, it's possible to get an approximate estimate of how far the instance is away by network distance:
$ httping -c20 -Qlg https://lemmy.world/setup|grep ^round
round-trip min/avg/max = 226.3/260.0/348.7 ms
$
226.3 milliseconds. How about lemmy.today, my home instance (which doesn't use CloudFlare, but is useful as a control)?
$ httping -c20 -Qrlg https://lemmy.today/setup|grep ^round
round-trip min/avg/max = 52.9/61.2/148.9 ms
$
52.9 milliseconds.
In lemmy.today's case, CloudFlare isn't in the picture, so we can check ICMP latency to get a feel for page generation time:
$ ping -c20 lemmy.today|grep ^rtt
rtt min/avg/max/mdev = 40.375/40.884/41.708/0.347 ms
$
40.375 milliseconds. Since we measured the response time for the webpage to be 52.9 milliseconds, figure that one might expect something on the order of 12 milliseconds for the page to be generated (maybe less than this; I don't know whether lemmy.today has other proxy infrastructure in place. One could maybe put tighter bounds on page generation latency by setting up a dedicated instance and measuring time for this to be returned).
So we can figure maybe something like 214 milliseconds to lemmy.world, about the time from US West Coast to Europe. Some other instance network distances from me:
$ httping -c20 -Qlg https://lemmy.zip/setup|grep ^round
round-trip min/avg/max = 211.0/220.2/283.9 ms
$
211 milliseconds. Also probably Europe (IIRC from reading their ToS at one point, they declare that they're subject to Finnish laws, so probably there).
sh.itjust.works says that it's Canada East Coast in their instance sidebar:
$ httping -c20 -Qlg https://sh.itjust.works/setup|grep ^round
round-trip min/avg/max = 117.9/133.0/192.2 ms
$
117.9 milliseconds.
Lemmy.ca (lemmy.ca does declare that it's in Canada already on lemmy.fediverse.observer):
$ httping -c20 -Qlg https://lemmy.ca/setup|grep ^round
round-trip min/avg/max = 42.7/49.3/158.2 ms
$
42.7 milliseconds, significantly closer than the sh.itjust.works Canada East Coast instance. So it'll probably be Canada West Coast.
For my purposes, that's good enough, since I really only care about whether an instance is near me or not. However, if one wanted to know where a distant instance is more than "not very near you" or "near you", one could extend this via "triangulating" an instance's network distance by measuring the latency from several locations; either via having several users in different locations participate and report latencies or setting up a VPS in several known locations around the world.