Weird web accesses from 104.244.79.113
I found a small series of odd HTTP requests from IPv4 address 104.244.79.113 while I was looking at Apache web server log files.
My Apache “combined” format log files run from 2025-03-13 to 2026-06-16 at the time of this writing. A machine at IPv4 address 104.244.79.113 made 2736 requests of my web site according to those log files.
| Date | Request count |
|---|---|
| 2026-01-19 | 2 |
| 2026-01-21 | 2 |
| 2026-05-02 | 2 |
| 2026-05-03 | 1 |
| 2026-05-06 | 2 |
| 2026-05-10 | 2 |
| 2026-05-11 | 2 |
| 2026-05-12 | 6 |
| 2026-05-21 | 1784 |
| 2026-06-06 | 933 |
Of the 2736 requests, 2630 were GET, 8 POST, and 98 apparently had no method at all. A typical combined format log file line for one of the puzzling 98 requests looks like this:
104.244.79.113 - - [06/Jun/2026:09:38:19 +0000] "-" 408 3335 "-" "-"
I’m not sure how you get a log line like that.
It looks like the (client) HTTP header was a single ‘-’ character.
The 3335 response bytes is similar to what I get when sending a
zero-length line to my Apache server with openssl s_client -connect 23.150.41.170:443
The 2026-06-06 requests all arrived 09:00Z, which is 11am in Luxembourg (more on that later). The 2026-05-21 requests arrived about half and half 03:00Z and 04:00Z, which is 5am and 6am in Luxembourg. I can’t read anything into time-of-day arrivals.
The smaller pairs or bursts of requests, occurring on 2026-01-19, 2026-01-21, 2026-05-02, 2026-05-03, 2026-05-06, 2026-05-10, 2026-05-11, 2026-05-12, are different than the 2026-05-21 and 2026-06-06 larger bursts. The appear to be probes for some webapps, notably OpenWrt’s luci management interface. These are the only POST method requests. Apache doesn’t save POST data in combined format log files, so I have no clue about what these mean.
What about 104.244.79.113?
Here’s what I can figure out from the IP address only.
whois and DNS
104.244.79.113 belongs to CIDR 104.244.72.0/21, which is currently allotted to BuyVM, Luxembourg. That address has no reverse-DNS lookup, SOA given as rdns1.frantech.ca.
traceroute
traceroute shows that two ways to route to that address.
On 2026-06-22, doing a traceroute from a VM located in Kansas City, MO, the routing is via Hurricane Electric. Another VM in Amazon’s Boardman, OR gave much the same trace.
traceroute to 104.244.79.113 (104.244.79.113), 30 hops max, 60 byte packets
1 lo.r3-er01x.mci1.us.rozint.net (23.150.40.5) 0.734 ms 0.646 ms 0.624 ms
2 hurricane-electric.gnd1.mci.kcix.net (206.51.7.5) 0.874 ms 1.028 ms 1.193 ms
3 * * *
4 port-channel13.core2.par2.he.net (184.104.197.10) 100.948 ms * *
5 be47.core3.par2.he.net (184.105.213.245) 102.414 ms 102.371 ms *
6 * * *
7 100ge0-57.core1.zrh3.he.net (184.104.193.133) 112.301 ms 111.404 ms 111.405 ms
8 10.255.240.10 (10.255.240.10) 112.597 ms 114.409 ms 114.327 ms
9 104.244.79.113 (104.244.79.113) 115.017 ms 115.559 ms 116.924 ms
From Highlands Ranch, CO, the routing is via Cogent.
6 * * *
7 be3758.ccr82.den01.atlas.cogentco.com (154.54.94.237) 8.325 ms 11.718 ms 11.700 ms
8 be8568.ccr32.oma02.atlas.cogentco.com (154.54.95.110) 31.010 ms be6904.ccr31.oma02.atlas.cogentco.com (154.54.95.98) 29.926 ms 29.914 ms
9 be5068.ccr42.ord01.atlas.cogentco.com (154.54.166.74) 29.902 ms be5214.ccr41.ord01.atlas.cogentco.com (154.54.165.134) 29.860 ms be5068.ccr42.ord01.atlas.cogentco.com (154.54.166.74) 29.850 ms
10 port-channel2718.ccr92.cle04.atlas.cogentco.com (154.54.7.130) 35.375 ms * 33.940 ms
11 port-channel3737.ccr91.dca04.atlas.cogentco.com (154.54.161.174) 49.453 ms be4986.ccr42.jfk02.atlas.cogentco.com (154.54.162.170) 116.000 ms port-channel3737.ccr91.dca04.atlas.cogentco.com (154.54.161.174) 48.881 ms
12 be2261.ccr41.par01.atlas.cogentco.com (154.54.47.166) 126.461 ms be3627.ccr41.par01.atlas.cogentco.com (66.28.4.198) 115.925 ms be2261.ccr41.par01.atlas.cogentco.com (154.54.47.166) 126.434 ms
13 be3123.rcr21.lys01.atlas.cogentco.com (154.54.36.166) 128.203 ms 126.881 ms port-channel3467.ccr92.lhr01.atlas.cogentco.com (154.54.94.46) 115.127 ms
14 be2830.rcr71.gva01.atlas.cogentco.com (154.54.57.34) 140.126 ms 139.421 ms 169.699 ms
15 be2983.rcr71.brn02.atlas.cogentco.com (154.54.57.46) 169.661 ms 169.642 ms 169.628 ms
16 be2830.rcr71.gva01.atlas.cogentco.com (154.54.57.34) 169.611 ms 149.11.48.17 (149.11.48.17) 169.596 ms 169.581 ms
17 104.244.79.113 (104.244.79.113) 132.435 ms 132.391 ms 143.061 ms
IP lookup services
iplocation.net says it’s owned by BuyVM in Luxembourg
geoiplookup.io says it’s in city of Bissen, Luxembourg
buyvm.net offers cheap VPS rentals, and one of the locations on offer is Luxembourg.
I’ll reluctantly believe that 104.244.79.113 is a BuyVM virtual machine in a data center in Luxembourg. Incidentally, Trustpilot has a 4.0 rating of BuyVM, with 16% (!) 1-star reviews. A lot of the 1-start reviews claim that BuyVM has a reputation for hosting sleaze and cybercrime-adjacent stuff.
Observations based on the log file
Looking at the log file lines created by requests from 104.244.79.113,
I see 104.244.79.113 requesting / on 21-May-2026T03:10:09Z,
without a referrer, and receiving a 301 status code.
The immediate next request is for / in the same second of the day,
referrer of “http://23.150.41.170/”.
The bot running on 104.244.79.113 uses HTTP to make a request.
My Apache is configured to redirect (via a 301 Moved Permanently status)
to https schema URLs.
For this request, both clear text and TLS requests give a user agent of “DuckDuckBot/1.1”,
which is a lie.
This is a recurring pattern: a request that gets a 301 from Apache,
eventually followed by a request for the same URL
with a referrer of “http://23.150.41.170/”.
This pattern repeats.
It’s a big clue because there’s lots of 301
redirection responses in my Apache logs.
Looking just at requests that get a 301 response
leads to a much clearer picture of how this bot operates.
The 104.244.79.113 bot makes a series of
HTTP requests of all the hyperlinks on my web site’s initial page.
That is, the bot is parsing the HTML
returned from it’s second request for / to decide what to get next.
The 104.244.79.113 bot asked for 389 unique blog posts in total.
My blog had 439 total blog posts May 21st,
so 104.244.79.113 bot got about 87% of the posts.
This is blog is generated by hugo.
The posts can be reached by “pages” of 10 post-summaries,
in reverse chronological order,
by Next and Previous links threading all the posts
together, and by topic pages, which hugo helpfully calls “tags”.
Additionally, there’s an “Archive” page which lists
all posts in reverse chronological order,
with monthly breaks to allow humans to keep their place.
The log file clearly shows the bot asking for every URL
it can parse out of the HTML it gets for requesting /,
starting with the menu bar URLs at the top of the page,
proceeding through the first 10 posts,
finally requesting /page/2/,
and then marching through the top level “tags”.
It asks for the URLs of the latest 11 posts in reverse chronological order,
suggesting that it got the posts’ URLs from the front page.
It probably didn’t get URLs from the archive menu bar item.
because after the latest 11 posts, it starts retrieving them in reverse
chronological order by tag.
It didn’t get the posts’ URLs from the Posts
menu bar item, because that leads to various /page/ URLs,
of which the bot only requested /page/2/.
After the first 10 posts on the first “autobiography” page,
it starts requesting posts from the first “books” tag page.
The bot makes requests for the images a post’s HTML references
before it requests the next post,
which is more confirmation that the bot parses HTML for further hyperlinks.
In addition to images, the bot also makes requests for index.xml,
Atom syndication format files, which are referenced
in /tags/*/ HTML.
Apparently the bot has some depth or number of links per “level” pruning.
I’ve categorized 50 posts under the “books” tag,
which means there are 5 “page” URLs under
the “books” tag.
The bot only asks for posts that appear on the first page of
the “books” tag index.
The bot only asks for /tags/books/page/2/ and /tags/books/page/3/,
but does not ask for posts from those pages.
Most of my posts have multiple tags, and thus appear as hyperlinks
in many /tags/*/index.html files.
There are very few (if any) requests that happen more than twice
in the May 21 burst, once getting a 301 and another getting a 200
response.
This means the bot has some kind of duplicate URL detection.
User Agent
The 104.244.79.113 bot used 15 different User Agent strings.
| User agent | Count |
|---|---|
| - | 101 |
| CCBot/2.0 (https://commoncrawl.org/faq/) | 183 |
| Mozilla/5.0 | 8 |
| Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ChatGPT-User/1.0; +https://openai.com/bot) | 219 |
| Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.3; +https://openai.com/gptbot) | 186 |
| Mozilla/5.0 (compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot) | 227 |
| Mozilla/5.0 (compatible; Applebot/0.1; +http://www.apple.com/go/applebot) | 185 |
| Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html) | 181 |
| Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) | 236 |
| Mozilla/5.0 (compatible; DuckDuckBot/1.1; +http://duckduckgo.com/duckduckbot.html) | 223 |
| Mozilla/5.0 (compatible; Google-CloudVertexBot; +https://cloud.google.com/vertex-ai-bot) | 198 |
| Mozilla/5.0 (compatible; GoogleOther; +https://developers.google.com/search/) | 203 |
| Mozilla/5.0 (compatible; OAI-SearchBot/1.3; +https://openai.com/searchbot) | 190 |
| Mozilla/5.0 (compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot) | 176 |
| Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots) | 219 |
| Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/135.0.0.0 Safari/537.36 | 1 |
The user agents appear to be chosen randomly, with the exception that the user agent for an initial 301 (HTTP) response is the same as the user agent for the follow-up 200 (HTTPS) response.
Some of these (ChatGPT-User, Amazonbot, Baiduspider and others) cause my Apache web server to run the infinite fake website generator. The bot dutifully parsed that HTML, and asked for hyperlinked files from within it.
Summary of Observations
- 104.244.79.113 ran an HTTP-first scanner. recycling requests that got a 301 for HTTPS.
- The scanner performs very mannerly HTTP, filling in
Referer:HTTPS headers when getting a 301. It doesn’t immediately re-request a URL via HTTPS, suggesting some kind of queue of URLs. - The scanner starts at
/, and works its way through hyperlink URLs that it can parse out of HTML given back for/. - The scanner asks for images as part of its crawling, in the order in which it encounters hyperlinks to them.
- The scanner has a depth or numerical limit on number of hyperlinked URLs
it will visit after parsing a directory’s
index.htmlfile. - The bot uses randomly chosen, brand name™ user agent strings.
Conclusion
Conclusion A
From one standpoint, it looks like this bot was crawling my entire website,
maybe for mirroring purposes, maybe for data to sell to an AI lab.
It did a methodical search of the site.
Conclusion B
From another standpoint, it was trying different user agents, mostly
AI frontier lab user agents, so it might have been testing the waters
for the real AI frontier labs to do something.
Both of these conclusions have their merits. The bot did methodically crawl 85% of my web site. It didn’t methodically try different user agent strings on a single blog post.
But it also only crawled 85%. It did not attempt a complete crawl, ignoring “page 2” and other ways to get all the URLs of my site.
I’m really torn about what to conclude, but this does illustrate that lots of inexplicable stuff happens on the internet.
Tools
I found my combined format log file parser to be of help. I used it to extract fields of log lines with particular properties, like all the URLs that were given a 301 response code. I probably could have done all the same work with some shell scripts, but it seemed easier with this dedicated program.
Speaking of shell scripts,
regular old pipelines helped me the most in this effort.
I did a lot of counting and summarizing using combinations
of grep, the combined format parser, sort, uniq -c and wc -l.
combined -f code logfile | sort | uniq -c
The above pipeline extracted all HTTP response codes from a log file and gave me a count of each code that appeared.