DB

D. Banta

info

Please Note

2 records found

An Approach for the Removal of Text and Ink Artifacts from Historical Watermark Images

Watermarks have an essential role in identifying the origins and age of specific documents. However, this is often a laborious process. One of the main issues in automatic watermark segmentation is the presence of text that obstructs it, making it difficult to properly reconstruct a watermark. Image processing and machine learning techniques face limitations, requiring time, training of data, or manual parameter selection. This research introduces a new method using wavelets transform to locate and remove text from a watermarked image, while preserving the underlying watermark. This method manages to outperform classic image processing techniques for the case when text is thicker than the watermark outline. ...
Watermarks are historical motifs present in the texture of paper that are commonly used to identify the paper manufacturers. They only become visible when viewed under certain light conditions. Under ideal circumstances, researchers may use watermarks to determine a historical document’s origins and context. To identify a watermark, it is matched to a previously archived watermark. Currently, this matching must be done manually, which is neither scalable nor parallelizable. Existing studies explore digital reconstructions of watermarks, but do not focus on a comparison-based setup. This report discusses a system that can automatically identify similar watermarks using traditional image processing techniques. The resulting system speeds up the process considerably, can be used on small datasets, and is more accessible to end-users.

The system uses harmonization, feature extraction, and similarity matching. Harmonization involves improving the clarity of the watermark, which is often obscured by the material properties of the paper. Feature extraction involves finding useful information from the isolated watermarks, and similarity matching uses this information to score the similarity of a pair.

We evaluated our system based on a dataset provided by the German Museum of Books and Writing. Over a broader range of quality, accuracy was found to be within the range of 41-53%. It was also found that improving watermark quality within the dataset improved accuracy results to around 82%. The system shows promise particularly with higher quality datasets. This report therefore demonstrates that traditional image processing techniques can be valuable when applied to situations where artificial intelligence may not be possible or efficient. Further research into this domain would be required to understand the advantages and limitations of image processing in comparison with artificial intelligence.
...