Weird web accesses from 104.244.79.113

I found a small series of odd HTTP requests from IPv4 address 104.244.79.113 while I was looking at Apache web server log files.

My Apache “combined” format log files run from 2025-03-13 to 2026-06-16 at the time of this writing. A machine at IPv4 address 104.244.79.113 made 2736 requests of my web site according to those log files.

Date Request count
2026-01-19 2
2026-01-21 2
2026-05-02 2
2026-05-03 1
2026-05-06 2
2026-05-10 2
2026-05-11 2
2026-05-12 6
2026-05-21 1784
2026-06-06 933

Of the 2736 requests, 2630 were GET, 8 POST, and 98 apparently had no method at all. A typical combined format log file line for one of the puzzling 98 requests looks like this:

104.244.79.113 - - [06/Jun/2026:09:38:19 +0000] "-" 408 3335 "-" "-"

I’m not sure how you get a log line like that. It looks like the (client) HTTP header was a single ‘-’ character. The 3335 response bytes is similar to what I get when sending a zero-length line to my Apache server with openssl s_client -connect 23.150.41.170:443

The 2026-06-06 requests all arrived 09:00Z, which is 11am in Luxembourg (more on that later). The 2026-05-21 requests arrived about half and half 03:00Z and 04:00Z, which is 5am and 6am in Luxembourg. I can’t read anything into time-of-day arrivals.

The smaller pairs or bursts of requests, occurring on 2026-01-19, 2026-01-21, 2026-05-02, 2026-05-03, 2026-05-06, 2026-05-10, 2026-05-11, 2026-05-12, are different than the 2026-05-21 and 2026-06-06 larger bursts. The appear to be probes for some webapps, notably OpenWrt’s luci management interface. These are the only POST method requests. Apache doesn’t save POST data in combined format log files, so I have no clue about what these mean.

What about 104.244.79.113?

Here’s what I can figure out from the IP address only.

whois and DNS

104.244.79.113 belongs to CIDR 104.244.72.0/21, which is currently allotted to BuyVM, Luxembourg. That address has no reverse-DNS lookup, SOA given as rdns1.frantech.ca.

traceroute

traceroute shows that two ways to route to that address.

On 2026-06-22, doing a traceroute from a VM located in Kansas City, MO, the routing is via Hurricane Electric. Another VM in Amazon’s Boardman, OR gave much the same trace.

traceroute to 104.244.79.113 (104.244.79.113), 30 hops max, 60 byte packets
 1  lo.r3-er01x.mci1.us.rozint.net (23.150.40.5)  0.734 ms  0.646 ms  0.624 ms
 2  hurricane-electric.gnd1.mci.kcix.net (206.51.7.5)  0.874 ms  1.028 ms  1.193 ms
 3  * * *
 4  port-channel13.core2.par2.he.net (184.104.197.10)  100.948 ms * *
 5  be47.core3.par2.he.net (184.105.213.245)  102.414 ms  102.371 ms *
 6  * * *
 7  100ge0-57.core1.zrh3.he.net (184.104.193.133)  112.301 ms  111.404 ms  111.405 ms
 8  10.255.240.10 (10.255.240.10)  112.597 ms  114.409 ms  114.327 ms
 9  104.244.79.113 (104.244.79.113)  115.017 ms  115.559 ms  116.924 ms

From Highlands Ranch, CO, the routing is via Cogent.

 6  * * *
 7  be3758.ccr82.den01.atlas.cogentco.com (154.54.94.237)  8.325 ms  11.718 ms  11.700 ms
 8  be8568.ccr32.oma02.atlas.cogentco.com (154.54.95.110)  31.010 ms be6904.ccr31.oma02.atlas.cogentco.com (154.54.95.98)  29.926 ms  29.914 ms
 9  be5068.ccr42.ord01.atlas.cogentco.com (154.54.166.74)  29.902 ms be5214.ccr41.ord01.atlas.cogentco.com (154.54.165.134)  29.860 ms be5068.ccr42.ord01.atlas.cogentco.com (154.54.166.74)  29.850 ms
10  port-channel2718.ccr92.cle04.atlas.cogentco.com (154.54.7.130)  35.375 ms *  33.940 ms
11  port-channel3737.ccr91.dca04.atlas.cogentco.com (154.54.161.174)  49.453 ms be4986.ccr42.jfk02.atlas.cogentco.com (154.54.162.170)  116.000 ms port-channel3737.ccr91.dca04.atlas.cogentco.com (154.54.161.174)  48.881 ms
12  be2261.ccr41.par01.atlas.cogentco.com (154.54.47.166)  126.461 ms be3627.ccr41.par01.atlas.cogentco.com (66.28.4.198)  115.925 ms be2261.ccr41.par01.atlas.cogentco.com (154.54.47.166)  126.434 ms
13  be3123.rcr21.lys01.atlas.cogentco.com (154.54.36.166)  128.203 ms  126.881 ms port-channel3467.ccr92.lhr01.atlas.cogentco.com (154.54.94.46)  115.127 ms
14  be2830.rcr71.gva01.atlas.cogentco.com (154.54.57.34)  140.126 ms  139.421 ms  169.699 ms
15  be2983.rcr71.brn02.atlas.cogentco.com (154.54.57.46)  169.661 ms  169.642 ms  169.628 ms
16  be2830.rcr71.gva01.atlas.cogentco.com (154.54.57.34)  169.611 ms 149.11.48.17 (149.11.48.17)  169.596 ms  169.581 ms
17  104.244.79.113 (104.244.79.113)  132.435 ms  132.391 ms  143.061 ms

IP lookup services

iplocation.net says it’s owned by BuyVM in Luxembourg

geoiplookup.io says it’s in city of Bissen, Luxembourg

buyvm.net offers cheap VPS rentals, and one of the locations on offer is Luxembourg.

I’ll reluctantly believe that 104.244.79.113 is a BuyVM virtual machine in a data center in Luxembourg. Incidentally, Trustpilot has a 4.0 rating of BuyVM, with 16% (!) 1-star reviews. A lot of the 1-start reviews claim that BuyVM has a reputation for hosting sleaze and cybercrime-adjacent stuff.

Observations based on the log file

Looking at the log file lines created by requests from 104.244.79.113, I see 104.244.79.113 requesting / on 21-May-2026T03:10:09Z, without a referrer, and receiving a 301 status code. The immediate next request is for / in the same second of the day, referrer of “http://23.150.41.170/”. The bot running on 104.244.79.113 uses HTTP to make a request. My Apache is configured to redirect (via a 301 Moved Permanently status) to https schema URLs. For this request, both clear text and TLS requests give a user agent of “DuckDuckBot/1.1”, which is a lie. This is a recurring pattern: a request that gets a 301 from Apache, eventually followed by a request for the same URL with a referrer of “http://23.150.41.170/”. This pattern repeats. It’s a big clue because there’s lots of 301 redirection responses in my Apache logs.

Looking just at requests that get a 301 response leads to a much clearer picture of how this bot operates. The 104.244.79.113 bot makes a series of HTTP requests of all the hyperlinks on my web site’s initial page. That is, the bot is parsing the HTML returned from it’s second request for / to decide what to get next.

The 104.244.79.113 bot asked for 389 unique blog posts in total. My blog had 439 total blog posts May 21st, so 104.244.79.113 bot got about 87% of the posts. This is blog is generated by hugo. The posts can be reached by “pages” of 10 post-summaries, in reverse chronological order, by Next and Previous links threading all the posts together, and by topic pages, which hugo helpfully calls “tags”. Additionally, there’s an “Archive” page which lists all posts in reverse chronological order, with monthly breaks to allow humans to keep their place.

The log file clearly shows the bot asking for every URL it can parse out of the HTML it gets for requesting /, starting with the menu bar URLs at the top of the page, proceeding through the first 10 posts, finally requesting /page/2/, and then marching through the top level “tags”. It asks for the URLs of the latest 11 posts in reverse chronological order, suggesting that it got the posts’ URLs from the front page. It probably didn’t get URLs from the archive menu bar item. because after the latest 11 posts, it starts retrieving them in reverse chronological order by tag. It didn’t get the posts’ URLs from the Posts menu bar item, because that leads to various /page/ URLs, of which the bot only requested /page/2/. After the first 10 posts on the first “autobiography” page, it starts requesting posts from the first “books” tag page. The bot makes requests for the images a post’s HTML references before it requests the next post, which is more confirmation that the bot parses HTML for further hyperlinks. In addition to images, the bot also makes requests for index.xml, Atom syndication format files, which are referenced in /tags/*/ HTML.

Apparently the bot has some depth or number of links per “level” pruning. I’ve categorized 50 posts under the “books” tag, which means there are 5 “page” URLs under the “books” tag. The bot only asks for posts that appear on the first page of the “books” tag index. The bot only asks for /tags/books/page/2/ and /tags/books/page/3/, but does not ask for posts from those pages.

Most of my posts have multiple tags, and thus appear as hyperlinks in many /tags/*/index.html files. There are very few (if any) requests that happen more than twice in the May 21 burst, once getting a 301 and another getting a 200 response. This means the bot has some kind of duplicate URL detection.

User Agent

The 104.244.79.113 bot used 15 different User Agent strings.

User agent Count
- 101
CCBot/2.0 (https://commoncrawl.org/faq/) 183
Mozilla/5.0 8
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ChatGPT-User/1.0; +https://openai.com/bot) 219
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.3; +https://openai.com/gptbot) 186
Mozilla/5.0 (compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot) 227
Mozilla/5.0 (compatible; Applebot/0.1; +http://www.apple.com/go/applebot) 185
Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html) 181
Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) 236
Mozilla/5.0 (compatible; DuckDuckBot/1.1; +http://duckduckgo.com/duckduckbot.html) 223
Mozilla/5.0 (compatible; Google-CloudVertexBot; +https://cloud.google.com/vertex-ai-bot) 198
Mozilla/5.0 (compatible; GoogleOther; +https://developers.google.com/search/) 203
Mozilla/5.0 (compatible; OAI-SearchBot/1.3; +https://openai.com/searchbot) 190
Mozilla/5.0 (compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot) 176
Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots) 219
Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/135.0.0.0 Safari/537.36 1

The user agents appear to be chosen randomly, with the exception that the user agent for an initial 301 (HTTP) response is the same as the user agent for the follow-up 200 (HTTPS) response.

Some of these (ChatGPT-User, Amazonbot, Baiduspider and others) cause my Apache web server to run the infinite fake website generator. The bot dutifully parsed that HTML, and asked for hyperlinked files from within it.

Summary of Observations

  1. 104.244.79.113 ran an HTTP-first scanner. recycling requests that got a 301 for HTTPS.
  2. The scanner performs very mannerly HTTP, filling in Referer: HTTPS headers when getting a 301. It doesn’t immediately re-request a URL via HTTPS, suggesting some kind of queue of URLs.
  3. The scanner starts at /, and works its way through hyperlink URLs that it can parse out of HTML given back for /.
  4. The scanner asks for images as part of its crawling, in the order in which it encounters hyperlinks to them.
  5. The scanner has a depth or numerical limit on number of hyperlinked URLs it will visit after parsing a directory’s index.html file.
  6. The bot uses randomly chosen, brand name™ user agent strings.

Conclusion

Conclusion A
From one standpoint, it looks like this bot was crawling my entire website, maybe for mirroring purposes, maybe for data to sell to an AI lab. It did a methodical search of the site.

Conclusion B
From another standpoint, it was trying different user agents, mostly AI frontier lab user agents, so it might have been testing the waters for the real AI frontier labs to do something.

Both of these conclusions have their merits. The bot did methodically crawl 85% of my web site. It didn’t methodically try different user agent strings on a single blog post.

But it also only crawled 85%. It did not attempt a complete crawl, ignoring “page 2” and other ways to get all the URLs of my site.

I’m really torn about what to conclude, but this does illustrate that lots of inexplicable stuff happens on the internet.

Tools

I found my combined format log file parser to be of help. I used it to extract fields of log lines with particular properties, like all the URLs that were given a 301 response code. I probably could have done all the same work with some shell scripts, but it seemed easier with this dedicated program.

Speaking of shell scripts, regular old pipelines helped me the most in this effort. I did a lot of counting and summarizing using combinations of grep, the combined format parser, sort, uniq -c and wc -l.

combined -f code logfile | sort | uniq -c

The above pipeline extracted all HTTP response codes from a log file and gave me a count of each code that appeared.