why html? just saving the json it fetches in the first place would be a lot easier to work with IMO, but it depends on what you're doing with it
Web Development
Welcome to the web development community! This is a place to post, discuss, get help about, etc. anything related to web development
What is web development?
Web development is the process of creating websites or web applications
Rules/Guidelines
- Follow the programming.dev site rules
- Keep content related to web development
- If what you're posting relates to one of the related communities, crosspost it into there to help them grow
- If youre posting an article older than two years put the year it was made in brackets after the title
Related Communities
- !html@programming.dev
- !css@programming.dev
- !uiux@programming.dev
- !a11y@programming.dev
- !react@programming.dev
- !vuejs@programming.dev
- !webassembly@programming.dev
- !javascript@programming.dev
- !typescript@programming.dev
- !nodejs@programming.dev
- !astro@programming.dev
- !angular@programming.dev
- !tauri@programming.dev
- !sveltejs@programming.dev
- !pwa@programming.dev
Wormhole
Some webdev blogs
Not sure what to post in here? Want some web development related things to read?
Heres a couple blogs that have web development related content
- https://frontendfoc.us/ - [RSS]
- https://wesbos.com/blog
- https://davidwalsh.name/ - [RSS]
- https://www.nngroup.com/articles/
- https://sia.codes/posts/ - [RSS]
- https://www.smashingmagazine.com/ - [RSS]
- https://www.bennadel.com/ - [RSS]
- https://web.dev/ - [RSS]
I second that - can't you just use the API and request the post and its comments as json?
Unless you want to explicitly archive the post as a human viewable HTML, the structured json from the API is almost always superior to further analyze, archive,... the posts.
Archiving thwmin a human readable form is exactly what I want to do.
Then I'm not sure if curl is up to the task in this case - most modern websites are PWAs that dynamically request data from their Backen and display it. Curl just fetches the initial HTML but doesn't execute the JS and thus might not see the content of the page as it isn't there yet the moment it gets saved.
Some PWAs "render" the first page contents in your HTML so that it is included and do not need to be requested.
Maybe try the alternative frontends if some of them work better? Also remember that feddit.org and many other instances deploy Anubis and might lock you out from making simple requests with curl
Got it working with wget
Fwiw there's SinglePage extension for web browsers. Maybe you can get a headless way to run that... Which frontend would you save from? I'd say i rather prefer to get json data from backend
I am preferring the standard Lemmy UI, but I already got it working.
Your browser dev tools network tab can be set to filter by CSS files. Then you just copy those URLs, and sort out the unnecessary garbage if there is some (ad frameworks, etc.). Since CSS won't really be changing a lot, you can just download it once and include it statically using your script and some <link rel="stylesheet" href="./path/to/file.css"> tags in the <head>
Mhm you could run a script through the HTML and look for stylesheet hrefs and download next to them.
I was more looking at doing this with a tool like wget. I got the following command, that downloads the post and its assets, but for some reason it then does not actually work:
'wget --recursive --no-clobber --page-requisites --html-extension --convert-links --restrict-file-names=windows --domains "feddit.org" --no-parent "feddit.org/post" "https://feddit.org/post/33942692"'
The advantage of wget is, that it also accepts multiple links at once, so it is quite easy do download multiple posts at once (which is my goal)
If this is what you are seeing that makes it "not work":

It seems to be implemented in a script that runs when the saved page loads and hides the content, even though that content is actually there. You can tell wget to ignore scripts (which aren't very useful in a download anyway, as they mostly pertain to interactivity) with --reject js. I tried this and was able to view the downloaded post properly.
I've no idea why the script does that, one could probably look at the lemmy-ui source code to figure it out.
That was my exact problem. Thanks for the solution.