Solving It On the Front-end
Requirements
As an input we have a website with some PDFs (documentation) for various modules and versions of some software in German and in English. We actually don't know:
- If all the docs exist for the latest version
- If the docs are duplicated in the different folders using different filenames
- If the docs are available only in one language or in both
As I'll be the one using these docs, I thought of an HTML page with an integrated search where there's a table with the following fields:
- Doc type
- Doc name (the actual title, not the filename)
- Link to the English version
- Link to the German version
That sounds reasonable, right? Let's get into it then.
The tools
The idea is to export all the information into Excel and do some matching there, then identify the duplicates, delete them, and finally create the resulting page.
The actual process involved choosing the tools along the way, but at the end I've got this:
- Chrome extension to scan the page and download all PDFs
- Grid.js for the search
- Pico CSS, because I like the nice and accessible look (it's used on this website as well)
- Some custom JavaScript to make Pico CSS styles work for Grid.js
- Some custom JavaScript to output the link to the file in the EN and DE columns
- Some CLI commands like
lsandfindto create the list of the files pdfinfoto extract the "Title" metadata from PDFs, if available (combined with some bash scripting)xargswithrmto remove the files from the list- A few
sedcommands to clean the CSV file saved in Excel from some Windows-specific symbols csv2list, a custom Python script to convert the result in Excel into the "list of lists" format required by Grid.jslist2csv, a custom Python script for conversion backwards- Excel for basic VLOOKUP operations and manual fixes
The sequence
I had to do some back-and-forth movements, so it took me about a day to make 300 lines from 600 files, but for the future I would do something like this:
- Download all the files
- Create a CSV with filename + folder and name (title) columns
- Match the files (duplicates, newer version, EN-DE) in Excel
- Create in Excel the list of files to remove and save to txt
- Remove the files with
xargs+ txt - Save the table from Excel to CSV
- Convert CSV to list of lists
It works quite well, and combined with fast Chrome rendering and an ability to open PDFs directly in the browser, it looks like a solid solution.
Conclusions
I'm not a programmer, so I asked an AI chatbot to do the heavy lifting with the programming. You still need to understand the Python or JS code it produces, and bash scripts are particularly dangerous. Looking back, I'd just move slower and more systematically to eliminate manual work and associated mistakes as much as possible. Well, you can't know everything in advance. I thought it was a 2-hour task, but I'm still happy with the result every time I use it.
The issue with the static HTML page is that it can work well only in a few cases:
- From a network folder, accessible to everyone, and when people are disciplined enough to open
index.html - When an HTTP server delivers it, which requires some setup
For example, to integrate it into SharePoint you would need to create a component and install this component. Which is kind of sad, because I can't see how a static page with some JS can be dangerous.
- Previous
.net 10 Could Hit the Sweet Spot