brandleys

Content and copyright

Republication and scraping.

Somebody has taken your material and put it somewhere else, either by hand or by script. Sometimes it is a competitor rebuilding your site around their own contact details. Sometimes it is an operation harvesting everything you publish and feeding it into a product that competes with you. The routes available differ, and so does whether it is worth using them.

What this looks like when it goes wrong

The most obvious version is a site that is yours with the names changed. The pages, the copy, the images, the structure and sometimes the staff photographs and the testimonials, sitting on a different domain, trading in the same market and ranking against you in the same searches. It reads as clumsy rather than sophisticated, and it still takes customers.

The second is automated and impersonal. Listings, prices, descriptions, specifications, availability and reviews harvested continuously and used to populate a competing product. No individual took a decision to copy any particular item. A script runs, and what you spent years building is available to somebody else the moment you publish it.

The third is republication that arrives dressed as a favour. An aggregator or a syndicator reproduces articles in full, credits you, links back, and takes the readers and the advertising income with it. Permission was never asked for and attribution is being offered as though it were a substitute. Whether this is a problem at all depends entirely on whether the traffic returns.

The fourth is your material inside somebody else's product, where it no longer looks like copying. A comparison site, a database, an application, a newsletter or a research report, built substantially on what you publish. The people using it may not know where it came from, and the operator may sincerely believe that changing the format changed the position.

What actually decides it

Begin with what is actually protected, because scraped material is often a mixture. The written copy, the photographs, the design elements and the code are ordinarily capable of carrying copyright. Prices, specifications, addresses and other plain facts ordinarily are not. Where the value lies in the collection rather than in any individual entry, the relevant right may be a right in the database or in the compilation rather than in each item, and that has its own conditions and can sit with a different party. Which right is engaged decides who can act and what they can ask for.

Copying has to be shown rather than asserted. Identical wording, retained formatting quirks, and errors reproduced faithfully are the sort of thing that makes derivation obvious, and where only facts have been taken it can be considerably harder. This is also why what is published, and when, is worth recording as a matter of routine rather than reconstructing after somebody has taken it.

Copyright is not the only route and frequently not the fastest. Terms of use can address automated access and create a contractual position independent of any intellectual property right, and in some circumstances interference with a system engages obligations of a different kind again. Alongside that sit the intermediaries: the host, the registrar, the search engines that index the copy, the network in front of it and any advertising or payment services behind it. A great many republication problems are solved at that layer while correspondence is still being drafted.

Then there is what you actually want at the end. Removal, delisting, attribution on your terms, a licence, or simply that it stops competing with you in search. An operator who is willing to pay is a customer with poor manners, and some of these end in an arrangement rather than a claim. An operator who is anonymous, offshore and uninterested is a containment exercise. Deciding which of the two you are dealing with should happen before the first letter, not after the third.

What we do

Separate what is protected from what is not

Which parts of the copied material carry a right, and which right, before anything is asserted.

Evidence the copying

Capturing what was taken and what it was taken from, in a form that stands up later.

Removal and delisting

Working with hosts, registrars, search engines and the services in front of a site rather than only its owner.

Fix your own terms

Site terms and technical signals that state your position on automated access and republication.

Turn it into a licence

Where the better outcome is payment and controlled terms rather than removal.

Deal with a competitor

Where the copying is by somebody in your market, handled so it stops rather than moves.

When to spend nothing

Republication that sends you readers and does not compete with you is usually best left alone, and so are low quality mirrors that nobody finds. Chasing every copy of an article is an expensive way to feel busy. The versions worth acting on are the ones taking customers, the ones outranking you on your own material, and the ones building a product out of what you publish.

The cheapest work here is preventative and does not involve anybody else. Terms that address automated use, a publication record that shows what existed and when, and a clear internal view of which material is genuinely valuable. None of that stops a determined operator, but it converts a difficult argument into a short one, and it costs a fraction of what the argument costs.

Before anything is sent

Positions harden the moment the other side takes advice, and the quiet routes stop being available once a demand has gone out. While nothing has been sent, everything is still open.