With Nitter.net receiving a cease-and-desist from the dark side, it's time I try harder to get something useful from the back-up I made before leaving twitter ... If we do just
sed -e "s/<h1/\n<h1/g" Alltweetsfromblogger.html | \
grep 935254889919369216 | sed -e "s:</p>:</p>\n:g" | \
cl xml
we end up with two "conversation" blocks, because the way I proceeded was
- paste the link
- paste the copied contents
- "Conversation" with `<h1>` tag is something that comes from the copy-paste
That was the starting point.
- Agreed, the command line is ugly, and its output isn't miuch better.
- Agreed, its scope is fairly limited: Alltweetsfromblogger is a thing I crafted on Google Documents with a manual routine from my list of threads this blog has links to.
- There are still some problems after you extracted as one line the : both before-conversation and in-conversation links are undistinguishable, apart from the fact one points to a date, which is weak
* but * Google Docs also offers .md export, and looking at the .md conversion, I did spot blank lines while there are none in the conversation itself. So maybe adding a blank-line-to-newline converter before we look for the post ID will help ?
sed -e "s/<h1/\n<h1/g;\
s:<p class=.c0 c3.><span class=.c2.></span></p>:<hr/>\n:g" Alltweetsfromblogger.html |\
grep "935254889919369216[^<]" | sed -e "s:</p>:</p>\n:g" | cl xml
That was the start, somewhere quite near after nitter.net disappeared. The idea since then is to extend the "blogpress" scripts, and have the "development stories" as well as chats with fellow developers aside the main blog text.
I stuck to the html export: the .md file would have been easier to parse, but the embedded pictures in base64 encoding promised to be a pain to work with. It's not the best software in the world, but I managed to couple it with a few lines that scan the json files from my old twitter backup so that I can identify when pictures in the Google Document are actually thumbs of videos, and reinject the link to the videos.I did some planning-on-calendar that truly helped nailing down some details, but the photo I shot is completely unreadable :-/ Fuzzy blogging for fuzzy coding ...
It's been about 1 month now and finally I get the kind of output I want. It's time to let you know about it and see how I can actually get those "threads" embedded into the sample chapter.








Vote for your favourite post
