DSheet: Beyond the PDF


Engineering datasheets are a fundamental source of component information, but PDF is designed primarily as a document interchange and presentation format. For an engineering platform that needs to index, query, and reference information inside thousands of datasheets, I wanted more control over how these documents are stored and accessed.
I developed Dsheet, a purpose-built document format for engineering datasheets. The conversion pipeline extracts the document’s text, SVGs, images, fonts, and page structure and reconstructs them into a compact binary representation. Deduplicated images and data can be reduced to one refrence source and images can be transcoded into more efficient formats such as WebP.
Dsheet Viewer (V1)
Download Orginal PDF To Compare (700kiB)From the above Dsheet the original PDF file size is approximately 700 kB, while the converted DSheet file is approximately 48.82 kB. The converted DSheet is therefore approximately 14.3× smaller, corresponding to a file-size reduction of approximately 93%.
Here is a small table of some of the files I have tested. The Simple files is just text and no images or svgs. While the complex has all these elements with the Complex 1 being the above Dsheet.
| Test | PDF | .dsheet | Size retained | Reduction | | Simple #1 | 133.2 KiB | 2.4 KiB | 1.80% | 98.20% |
| Simple #2 | 116.4 KiB | 3.0 KiB | 2.58% | 97.42% |
| Complex #1 | 699.2 KiB | 48.9 KiB | 6.99% | 93.01% |
| Complex #2 | 1.3 MiB | 322.7 KiB | ~24.3% | ~75.7% |
As you can see, these are really good early results in terms of file-size reduction. I believe these numbers could be improved further, as the current binary format is essentially just a conversion of the original JSON test data. Changing the data to use a more efficient binary scheme should improve these numbers significantly, which would be part of a Version 2.Current Problems
As you can see in the preview, some of the text and SVGs have alignment issues. These problems come from how the text is grouped and how the SVG dimensions are reduced. Resolving this would require significantly more work on the input pipeline, but overall, I would conclude that the current results are approximately 96% of the way there.
Another area for improvement is the handling of layered images with masks. The current methods handle normal images well, but they do not appear to be compressing masked images as effectively, as seen in the top datasheet.
For example, if we take the top file, which is 48.9 KiB, the file is made up of images, pages, and SVGs. The image data alone accounts for approximately 17 KiB and is stored as WebP, making it the largest contributor to the file size. This should be addressable with another pass over the image-processing portion of the pipeline.
Future Improvements
The system is currently in a working state, but I do want to modify it in the future to fix the remaining layout bugs and further reduce file sizes.I believe that I could get the current document shown to a 97% reduction in size with better processing techniques.
The current system was designed as part of my engineering platform project. One of the main features I want to add next is provenance for the file structure. Since we are already generating a JSON file as part of the pipeline, this data can eventually be fed directly into the database to automatically populate component information.
This would also allow for much faster referencing of specific figures and sections within datasheets. For example, if a diode is specified as having a peak reverse voltage of 40 V, I could click on that value and have the datasheet open directly to the relevant location, with the corresponding section highlighted.
However, that functionality will be part of Version 2 of the system.