Skip Navigation

InitialsDiceBearhttps://github.com/dicebear/dicebearhttps://creativecommons.org/publicdomain/zero/1.0/„Initials” (https://github.com/dicebear/dicebear) by „DiceBear”, licensed under „CC0 1.0” (https://creativecommons.org/publicdomain/zero/1.0/)S
Posts
97
Comments
2511
Joined
3 yr. ago

  • The BS (trackers etc.) are on the web sites and you get them automatically unless your browser either actively filters them out, or implements an incomplete enough part of present-day web protocols that they don't receive the trackers (e.g. by not supporting Javascript).

    Unfortunately, filtering trackers is a never-ending arms race, while using a browser for the "historical web" makes a lot of present-day sites inaccessible.

    I.e. the problem is the web itself, not the browser, for the most part.

    I use Firefox but have fairly aggressive umatrix origin (ad blocker) settings, if that helps.

  • I had no idea that 2G worked til just now. Damn. I had to reluctantly mothball my 3G phone some years ago, but maybe there was a way to keep using it on 2G until now.

  • FFF won't work on that site right out of the box, because of the bot challenge. You'll need a workaround. FFF has the same problem with fanfiction.net (FFN) which is one of the biggest fanfic sites. So there are a number of FFF github issues and doc entries related to FFN and looking at those might help. I do know that FFN is scrapable using browser orchestration. https://github.com/FicHub/fichub.net may have some code for that, but getting your own instance running will be quite a bit more headache than just running fanficfare.

  • I tried wget and got a small index file that looks like a bot challenge, maybe Cloudflare Turnstile. I don't know of a simple automated workaround. But, Turnstile gives you a rewritten url that then goes back to the forum page and sets a cookie, iirc. So if you can visit the page with a real browser, then save the cookie and transfer it to wget (there's some option to set an arbitrary header) that's one thing to try.

    There's another hack used by fanfiction.net readers, where if you want to save a multi-page story, you can manually visit each page with a browser, i.e. click "next" again and again to load all the pages. Up to a few dozen such clicks isn't so bad. Fanficfare (fanfic downloading program, sometimes abbreviated FFF, https://github.com/JimmXinu/FanFicFare ) then has an option to retrieve the pages from your on-disk browser cache instead of trying to get them from the remote server. That's another approach you can try, either with fanficfare or your own scripts.

    The site you're looking at uses xenforo which is a very popular forum server program. The actual layout of xenforo sites varies, but fanficfare probably already recognizes something similar, so try using one of those interfaces. I think spacebattles.net (another fic site) uses xenforo and FFF supports it, so it might be a good start. You will have to modify FFF to recognize civfanatics instead of spacebattles. It will help to know or pick up some Python, but you shouldn't have to become an expert.

    Added: if you really want to automate your scrape, you will have to orchestrate a browser as mentioned earlier. IDK if there is code around to already do it. If you can program, it's not terribly hard to use Puppeteer or Selenium, but it will take some farting around to deal with the site layout and anti-bot stuff. It's not guaranteed to work right off the bat, but with enough determination you can do it, especially if your scrape volume is low and you can run it slowly. I'm not deeply involved in this stuff (like you, I just occasionally want to download something for personal use) but there are tons of webpages and articles by people (who I'd mostly consider evil) who do it at scale.

  • At least til recently, Waymo here in California was using modified Jaguars. IDK if they started out as Jaguar EV's or if they were ICE cars converted by Waymo. But they are everywhere in SF and also seen in other parts of the region.

  • You're trying to run a scraper, and unfortunately a lot of AI companies are doing the same thing on such a big scale that it creates a DOS attack. So lots of sites now have anti-scraping measures. I clicked the forum link and saw a brief interstitial that looked like bot protection.

    I'm unfamiliar with httrack but wget fails pretty often by the site just rejecting the user agent. Try "wget -Dfoo [url]" to save the response headers in the file "foo", if I remember it right. That will let you check if there is an error code. You could also examine the too-small html index that you got, to see if it has error messages inside.

    Sometimes curl works when wget fails. For both of them, there are CLI options to set the user agent to something different.

    Getting images from forums often requires you to have a login cookie in your client. It's simplest to log into the site and then paste the cookie into your scraping program or script.

    The next thing after wget/curl would be to write a scraping script (say with python urrlib) that can analyze the html a little as it goes. The thing after that would be script an actual browser, with puppeteer or selenium etc.

    Regarding the links pointing to your local disk, that's probably because they are relative links, like "href=./foobar.html". You can make them point to the remote server by inserting an HTML BASE tag into the top of the file, like

    <BASE HREF="https://whateveritwas.com/forum">

    or whatever. If it's just for one or two pages you can do that manually. Otherwise, modify your scraping script to insert it, before or after saving the output.

  • The video kept stopping but after a few retries I got through it. It would be helpful to have a transcript. Anyway, I wasn't impressed.

  • Emacs! Emacs Makes All Computing Simple.

  • Some friendlier ones. There was once an expensive camera brand called "Exakta" but Exakta owners affectionately called them "Approxima".

    I worked near a slightly fancy Chinese restaurant called Royal East but my coworkers (who went there often) liked to call it "Royal Eats".

    There was a nearby supermarket called Purity Supreme but people called it Poverty Supreme. That was waaaay before Whole Foods.

    Amherst, Massachusetts was sometimes anagrammed to "Hamster". I mean you could address letters that way and they'd get there.

  • 12 years.

  • Who and who?

  • Usually the password is different, yes. If you're unfamiliar with Linux system administration you might ask for more specific help. Root is the administrative account that lets you add and delete user accounts, access all the system's files, and that sort of thing. Ultimately though, it's like the joke about how many psychiatrists it takes to change a light bulb (just one, but the light bulb has to want to change). If you want to break through your time limiting software, you will find a way to do so. So you have to be willing to go along with it.

  • Yeah I think lots of bars are doing this now, so that info about unruly customers propagates between them. I don't think I've gone into a bar since before November 2019 (when Covid hit) and will probably keep staying out of them. Back then at least though, folks old enough to obviously be over 21 usually didn't get carded. I guess that has changed.

  • I've never noticed that. Is it something new? Maybe it's to get in the way of AI scraping (automatic transcription of the audio track, fed into LLM). Kind of a nonstop audio CAPTCHA.

  • This is about producing tiny amounts of electricity by radioactive decay from tritium. Products like that have been around for a while and they have some niche applications, but they're quite expensive and inefficient. They make a comparison with tritium exit signs but those by now are a pretty bad idea, compared to LED's and conventional batteries. The exit signs create a nontrivial disposal problem. They do still exist.

    They say gadget has 20 to 40 curies of tritium so by my math that's around 5 milliwatts of thermal power (idk their electric conversion efficiency). About what you get from (corrected) under 0.5 square centimeters of solar panel averaged over 24h, plus some battery or capacitor storage. Doesn't seem that competitive.

  • Buying that many cpus, gpus, ssd's, and ram is also slow and expensive nowadays, thanks to the same idiots that are building the data centers.

  • Well I use duckduckgo instead of google. Maybe that makes me a ducky-eyed chud, heh.

  • IMHO they are pretty useless once the novelty wears off. You use it a few times and put it in the closet. So even a cheap model will be BIFL.

  • Space, the final frontier @lemmy.ml

    Alien World Chemistry Found Inside Meteorite

    www.seti.org /news/alien-world-chemistry-found-inside-meteorite/
  • Spaceflight @sh.itjust.works

    Space junk debris cloud discovered in high-traffic orbit 'is a potential minefield' for the costliest satellites

    www.space.com /space-exploration/satellites/space-junk-debris-cloud-discovered-in-high-traffic-orbit-is-a-potential-minefield-for-the-costliest-satellites
  • covid @hexbear.net

    COVID Nimbus Variant Now Leads the U.S. as Cases Grow in 27 States and Emergency Visits Rise

    www.medicaldaily.com /covid-nimbus-variant-nb181-dominant-27-states-er-visits-july-2026-476165
  • ultralight @lemmy.world

    Anker 511 Nano 3 USB-C charger weighs 37 grams

    www.anker.com /products/a2147
  • Free and Open Source Software @beehaw.org

    Current SOTA in local FOSS speech to text?

  • flashlight @lemmy.world

    Is there an Anduril headlamp that accepts 18650 button tops?

  • flashlight @lemmy.world

    Some AA and AAA lights

  • flashlight @lemmy.world

    Wurkkos TS30S Pro, opinions?

  • Not The Onion @lemmy.world

    South Africa withdraws AI policy after it was found written by AI

    www.the-independent.com /tech/ai-policy-south-africa-withdraw-b2966866.html
  • Not The Onion @lemmy.world

    Someone may have maliciously caused a huge rodent invasion in California

    www.sfgate.com /bayarea/article/california-nutria-reintroduced-22198180.php
  • Privacy @lemmy.world

    FBI Extracts Suspect’s Deleted Signal Messages Saved in iPhone Notification Database

    www.404media.co /fbi-extracts-suspects-deleted-signal-messages-saved-in-iphone-notification-database-2/
  • flashlight @lemmy.world

    Cell welder advice?

  • Privacy @lemmy.world

    OkCupid gave 3 million dating-app photos to facial recognition firm, FTC says

    arstechnica.com /tech-policy/2026/03/okcupid-match-pay-no-fine-for-sharing-user-photos-with-facial-recognition-firm/
  • ultralight @lemmy.world

    Mal*Wart 2x2032 27 gram headlamp

  • flashlight @lemmy.world

    Mal*Wart 2x2032 27 gram headlamp

  • flashlight @lemmy.world

    Chemical lightstick lumen/runtime test request

  • ThinkPad @lemmy.ml

    T520 fan has stopped spinning

  • Privacy @lemmy.world

    Large-Scale Online Deanonymization with LLMs

    simonlermen.substack.com /p/large-scale-online-deanonymization
  • News @lemmy.world

    Amazon BUSTED for Widespread Scheme to Inflate Prices Across the Economy

    www.thebignewsletter.com /p/amazon-busted-for-widespread-price
  • Fuck AI @lemmy.world

    I want to wash my car. The car wash is 50 meters away. Should I walk or drive? | Stupid AI responses

    mastodon.world /@knowmadd/116072773118828295