Showing posts with label Bad SEO. Show all posts
Showing posts with label Bad SEO. Show all posts

Friday, 6 June 2008

How not to keyword spin

Unique content's a good thing, right? Not always.

To get well written text, you need to invest some time, or pay someone skilled to do it for you. Both are expensive and slow exercises, so many SEOs choose to take a shortcut and "spin" articles. This generally involves taking one source text, and altering the language inside it to "create" a new article. The better spinning programs are aware of things like grammar rules, word frequencies in various languages, set phrases and idioms, and so forth. The worse - and predominant - spinners simply perform synonym replacement, which produces this kind of mapping:

"We walked to the large house" would be replaced by "We ambulated to the gigantic residence".

The latter looks ugly, is awkward to read, and generally doesn't fulfil any kind of quality standard - though it is different, thus helping create unique content, and the meaning remains pretty much the same. The massive failing here, though, that utterly defeats the goal of those using spinners, is how trivial the product is to detect.

In every language, there are common words, rare words, and everything between. The probability of individual words occuring in a piece of text is fairly constant, being skewed a bit depending on the document's type and domain (agricultural reports will be more likely to contain terms about farming, horticulture, and plant and chemical proper nouns, for example). The core set of terms, and their frequencies, will remain the same.

The direct synonym replacement used by typical spinning programs (spinners) has two problems. Firstly, one must bear context in mind when picking alternative words. For example, "junk" (used as a noun) can also be:

"boat, clutter, debris, discard, dope, dreck, dump, flotsam, garbage, jetsam, jettison, litter, refuse, rubbish, salvage, scrap, ship, trash, waste"

Depending on whether we're using this word to talk about a sailing vessel or a piece of rubbish, we can divide this set of alternatives into two distinct groups with different meanings. Simply using a thesaurus to pick a random replacement word will often change the meaning of a sentence. "I thought your product was a heap of junk" is not semantically equivalent to "I thought your product was a heap of boat". Note how simplistic substitution also makes the sentence grammatically incorrect in this case.

The second problem with direct synonym replacement is that it doesn't care about the probability distribution of words in a language. This always leads to the inclusion of rarer words, and exclusion of more common ones. In one above example, we used ambulate instead of walk; the former is a comparatively rare word. Using it makes the sentence more awkward, and harder to read (protip: always use the simplest language that you can).

Just to show how easy spun pages are to spot, let's find one, and take it apart, then see how abnormal it is. We're going to first find how frequent words are in English, and then use them to compare a spun article to a previous post on my blog.

The reference frequency list we use to represent general English comes from the British National Corpus. This is in British English, so we'll make things fairer by Anglicising the spun document, making "color" into "colour", "center" into "centre", and "-ize" into "-ise".

We'll mathematically compare both the spun and un-spun text against this reference model of the English language. This can be done by first building a list of words used in a document, and then counting how many times this occurs in the document. Dividing the count of a word by the total number gives the probability than any random token from the document will match that word. We'll also have a list of these probabilities from the British National Corpus (BNC). To compare this, we'll take the absolute difference between measured and reference probabilities for each term, and express that as a percentage of the reference likelihood. As these percentages get pretty high, and to reduce the impact of any anomalous data, we'll also measure the log of the difference measure.


Spun document

Taken from Good Articles Recommend Top Rank by SEO, an almost illegible and probably spun document (it may possibly be badly translated by someone with a newfound love for thesaurii, though given the topic domain - SEO - this seems only minimally likely).

wordfreqprobbncprobdifferencelogdiff
seo50.01259.00E-09138888791.10%6.142667198
spell-check10.00251.80E-0813888789.11%5.142664384
overusing10.00251.90E-0813157794.88%5.119183112
copywriting20.0055.70E-088771829.78%4.943090196
scruffily10.00255.90E-084237188.04%4.627077738
overeat10.00258.80E-082840809.03%4.453442039
well-crafted10.00251.38E-071811494.22%4.258036952
proofread10.00253.05E-07819572.12%3.913587174

Full dataset

We can see a few words sticking out here. Some give an indication of the document's topic (SEO, copywriting) while others are quite bizarre (scruffily, overeat). The difference column shows the magnitude of frequency variation from what's expected - a difference of zero means that a word occurs just as frequently in this text as it does in the British National Corpus; a difference of 100% means that the word occurs twice as often or half as often. Note how the words that stick out hugely aren't that congruent as a set - overeating and scruffiness have little to do with copywriting, spell checking and proof reading.

Differences measured this way will be skewed rapidly by any rarer words that come into a document, and every document that has something to say will have to incorporate some topics using rarer that don't fit the curve perfectly - this would be expected. So, using this measure, a non-zero difference score is inevitable; logs have been taken to smooth differences in scale. What is significant is where and how much the differences are.

The mean difference from standard English for the unspun document is 947349.85%, and the mean of the logs of the difference measure is ~1.46. These numbers show us how different the words in the spun document are from what would be expected in general language.

Unspun document

Taken from my overall vaguely positive Ubercart review.

wordfreqprobbncprobdifferencelogdiff
cron20.0024449889.00E-0927166431.27%5.434032591
www20.0024449881.90E-0812868256.85%5.109519721
php10.0012224942.90E-084215396.07%4.624838387
firewall10.0012224943.80E-083216989.14%4.507449594
uploading10.0012224943.90E-083134499.74%4.496168238
todos10.0012224944.80E-082546762.24%4.405988403
upload10.0012224945.60E-082182924.83%4.33903878
metadata10.0012224945.80E-082107648.07%4.323798095
poin10.0012224945.80E-082107648.07%4.323798095

Full dataset

We can guess from the top differences here that the document is related to computing and fairly technical. The biggest differences in word probability are in the range of say 1e6 - 2.7e7, a lot less than the top four in the unspun document, which were from 8e6 all the way up to 1.4e8. The mean difference is half that of the unspun document (524475.55%) and the log differences again significantly smaller (1.18).

Comparison

For good measure, and to illustrate this point clearly, here's a graph. The red line is the spun document, the blue one the unspun one. For a document that completely followed average word frequency, you'd see a line at y=0 (i.e., a flatline).

Visual comparison of terms in a spun and unspun document

This shows that the spun document uses English consistently more unusually than the human-written (unspun) document; the red line is higher than the blue one, and the higher a point is, the more it varies from the British National Corpus' survey of English usage. For reference, that covers ~10 million words in 4000 documents, so it's a fairly good source of comparison data.

We all know about term frequencies (TF); it seems fair to guess that search engines have models of these, and that's it's not computationally intense for them to use TF as one tool to distinguish spam from useful content. When one can pick out spun content so easily (this system took ~40 minutes of coding and juggling in excel to make it look pretty, for one guy), there's really no point bothering to add it to your site.

Of course, a sophisticated document spinner is definitely possible to construct. My point here is, the cheap and common ones only provide a massive bright flag that your site is spam. Avoid them.

Data

A full set of all the produced data, in a pretty and readable format, including the full keyword data, and a large graph, is available online here. The texts actually used for comparison are here (unspun) and here (spun).

Further information

If you feel like exploring English word frequences and getting into that long tail, I can't recommend anything more highly than Wordcount.org.

Tuesday, 3 June 2008

No search engine traffic here please, and half you visitors can f*** off too

You gotta love sites that hate spiders as much as this. Open up IE (if you have it) and visit www.fdms.com. No problem, right? I mean, it's a pretty bad site, with some clear SEO problems, sucky design, lack of content, etc, but you can see something, at least (hopefully).

So, that's how things are meant to work - far from stunning, admittedly, but vaguely functional. Now, switch to something quite popular, say Firefox or Opera, and try the same. Oh! What's this?



IE4, you say.. that's what, a decade old this year? And I need a newer browser than that? Visitors aren't even given a chance to try their luck and saunter on in regardless. If developers are designing sites with IE4 as their target browser, something has to be seriously wrong. It fails with plenty of standards, meaning lots of exclusionist code bodges on the inside, and has had less than 0.5% market share with savvy users (see this list as W3Schools); even ignoring the bias that might be a list gathered there, IE4 hasn't shipped with any OS since Windows ME, 8 years ago. What the fuck are these guys on?

Anyway, luckily, hilarity ensues. Looking at the URL of this error page, we have http://www.fdms.com/ns_win.asp. Now, if they were detecting our browser using JavaScript, we'd still be able to see some kind of content somewhere - but we get HTTP-redirected to an error page. This looks like HTTP level user-agent sniffing. I wonder what the search engines think of that?

Free Image Hosting at www.ImageShack.us

Fantastic! We need Javascript and IE to browse this site, and if we're a spider, we can get lost. So what's the result?



Abject failure - 6 results, all of which save one are error pages and login links. Fantastic - keeping FDMS's developers where they belong - AS FAR AWAY FROM THE WEB AS POSSIBLE!

Thursday, 17 April 2008

BT Web Clicks = Bad

BT Web Clicks are sending out an AMAZING, unique offer!

They then proceeded to ask me “what do we rank for on google?”, my response was “your company name, unless you request otherwise”.

They then went on to mentioned that “the man from BT” can get us listed “at the top” of the search engine for “our keywords”.


Interesting! What does Google say about companies offering that?

No one can guarantee a #1 ranking on Google.


So.. BT (British Telecom) are officially offering a large pile of fail - and this is just the tip of the iceberg; inside are hidden £15k/month costs, outrageous consultancy fees, targetting at the meek and lonely, and to top it all, a cornucopia of grammatical errors. Watch out for this one, people - they may be cold calling in your area now.

James Wade has the full scoop: BT Web Clicks review.

Thursday, 10 April 2008

T-Mobile UK SEO Audit

Well, I've had this internal document on my desk for a few years, and it's kinda heavy - and certainly not pertinent enough any more to cause damage. So, here's a high-grade commercial SEO Audit, worth somewhere around 1500GBP at the time of production. Enjoy.

T-Mobile UK SEO Audit

I suppose the most shocking thing here is the quality of the site - thankfully it's improved since then - it's really incomprehensible that a company of this proportion failed so epically. Or, well, not.

You might like to take this audit and use it as a framework for your own; some things are well out of date, there are factual errors, and typos, but the formatting works, and the principles are still the same. Let me know what works for you!

Wednesday, 26 March 2008

AMEX advice for idiots

AMEX have been blasting out some fairly self-contradictory advice recently, according to WebProNews. They suggest that small businesses avoid employing SEO guys:

don’t waste money on so-called Search Engine Optimization (S.E.O.) specialists. Search engines are very quick to penalize sites that try to trick their filtering techniques, and once your site has been put on Google’s blacklist, it will take forever to get off.


Obviously the SEO people have been up in arms over this rather broad tarring. While of course there are many terrible SEO firms out there, and even more that over-charge, it's fairly easy for anyone to spot a dodgy outfit thanks to Google's published SEO guidelines.

Anyway, the bit where this gets "interesting" is about here, where AMEX then proceed to dish out SEO advice themselves!

using clean U.R.L.’s like yourdomain.com/store/widgets instead of yourdomain.com/store.php?id=42&categoryID= widgets will increase your chances of getting indexed in a search engine.


Sound advice indeed. Shame about the utterly conflicting points of view presented here, though. Tell you what, AMEX: how about you stick to raping me on chargeback costs, and I'll stick to what I do best, too.

Thursday, 19 July 2007

Optimaliser.no

Poor old Netty - they sell undertøy (lingerie / underwear). Scroll to the bottom of their front page; see a weird thing in the bottom right? Let your mouse cursor hover over it. Their "SEO" company has placed links to all their other clients on the frontpage! I say, sue the buggers. Don't use Optimaliser.no! They're selfish scum, abusing their clients. See video below for the full insult they've levied on their unfortunate customers.



If you don't get it, they've placed links on client's homepages that detract from the client's optimal setup, instead helping Optimaliser.no's business.

Friday, 6 July 2007

Custom PC

Take a look at this site.

http://www.kustompcs.co.uk/

It's really badly set up for SEO. Really, really badly. For example:


  • Who on earth searches for "kustom pc" instead of "custom pc"? Branding based on a misspelling is a big mistake.

  • If somebody does actually go for a "kustom" custom pc, surely they'd go for one - not many! so why would you buy kustompcs.co.uk? The targeted term here is "custom pc".

  • You can buy a custom PC from kustompcs.com or kustompcs.co.uk - great! The company's bought both domains. So why have they chosen to duplicate content between the sites, instead of set up a permanent redirect? Awful.

  • The title tag.. oh god, the title tag. This tag is critical. Once I placed an image-based form I was working on for a client on a test server, and used the name of the campaign for the title of the page. It happened that the form got spidered; when we went to review the progress of the client's site, our marketing guys found that my creative had stolen the number 1 spot for the name of the campaign, blowing the client's site out of the water. And what have these guys done with it? "Kustom PC's"? Not only can they not spell - the grammar's awful - but there aren't /any/ keywords in there! How about "Custom PC parts and builds at kustompcs.com". Easy. Child's play, in fact.


On top of this - no meta description, no h1, it's in tables (disrupting flow and stopping important content coming to the top), the left category nav uses javascript; problems, guys. No wonder you're so lowly ranked for Custom PC ;)

Friday, 29 June 2007

Yahoo! Stores - hard coded duplicate content

I read this post on the Yahoo! Stores blog. The Yahoo! stores blog is there to scratch the surface of online marketing for merchants new to the scene; it's probably great for giving people introduction to subjects, and leads to follow up, but for old dogs there's not a huge amount of new information. It certainly lets us see that Yahoo!'s helping its merchants, and that they're doing well from their help and the amazing Yahoo! store system. Anyway, in this SEO-oriented posts, Karl Ribas brought up some valid points, including a little intro to duplicate content:

Duplicate content was a pretty big concern at SMX, as having non-unique content on your website is quickly becoming a bigger and bigger problem for online merchants. ... From a search engine’s point-of-view, their one and only goal is to serve a variety of quality results per query, not multiple versions of the same content


Fantastic advice!

Yahoo! stores have cleverly helped us out here. As we all know, visiting the root URL - / - of a domain should really show the homepage; no redirects, no frames, just a plain and easy HTTP 200 response with some good content. And, to their credit, Yahoo! have managed this millimetre scale hurdle.

Now, we also know that in most circumstances, it's great to have a link to your homepage on every page in your site, right? After all, it's the most important page, and where people like to navigate from - so great to provide a link to in case they get lost.

Yahoo! have cottoned on to this little nugget of wisdom, and kindly added a link named "home" to the homepage of a site on every one of its sub-pages. Well, kind of. In fact, it's a hard-coded link, using the link text "home" (also hard-coded - heaven forbid anyone decides that using all-lower-case looks awful, or would prefer slightly less heterogenous link text here):

Dear Zack,

I think you might've got your wires crossed. I'd like to change the
small H is my yahoo stores / store editor system, I don't really mind
about yahoo web hosting. There are options to change all the other
tabs, but the name for the "home" page seems kind of elusive, even
though intuitively I expected them to be in the same place. Could you
check and come back to me ?


Hello Leon,

Thank you for contacting us.

It's not possible to change the 'H' in the navigation bar because the
links are hard-coded into the store software.

We apologize for the inconvenience.


This locked-down and widely shown link points to some strange, new page that's mentioned nowhere else in the store - to "/index.html".

<ul id="nav-general"><li><a href="index.html">home</a></li>


"What's this new-fangled index.html?" I hear you cry. "Where's my homepage?". Well, don't worry! Yahoo!'s kindly duplicated your homepage content for you onto this new URL. So search engines can NOT ONLY get your stuff at the root page, as standard, but now you'll find your link weight directly split between links to / - added by you - and links to /index.html - forcibly inserted by Yahoo!.

What do Yahoo! think of this? Can we get it changed?

Hello Leon,

Thank you for writing to Yahoo! Store Support.

Although this feature is not currently available in the Yahoo! Store
software, we do consider your feedback regarding the features you'd like
to see a very important part of how our development team decides which
features to add to the Yahoo! Store software.

We do not currently have an estimated time for if or when this feature
or any other features may be released. However, we do release a regular
newsletter to all of our merchants at the following link:

http://www.insightsforum.com/

You can see previous copies of the newsletters at:

http://store.yahoo.com/vw/merchant-newsletter.html

We appreciate your feedback. We've forwarded your comments to our
development team for review.

We believe this solution should resolve your issue, if it still
persists, please call us at 1-866-800-8092.

Please do not hesitate to reply if you need further assistance.

Regards,

Andre


Thanks Andre! I'm not sure what led you to believe it should resolve my issue, I'm fairly sure you just told me that it wasn't resolvable. Have you tried visiting http://www.insightsforum.com/ ? I'll save you the trouble:

Bad Request (Invalid Hostname)



Well, maybe the archive mentioned has something useful. Let's take a look at the last post:

February 2005-
Note: Insights has switched formats. While you will continue to receive monthly newsletters, all articles are archived on the Insights Forum site rather than a single HTML file.


Thanks Yahoo!. That's pretty good.

Will you stop duplicating my content soon please?

 
Marketing & SEO Blogs - Blog Top Sites sitemap