Sunday, February 15, 2009

Solution for Duplicate page listings - 'canonical' attribute in head tag

One sort of duplicate content happens because of plagiarism, copying to forums, copying of articles to blogs (see  How Google Treats Duplicate Content). In those cases, there are really several physical instances of a page for various reasons. 
 
But there is another sort of "duplicate content" that is often really just an artifact of how the Web works and in part a bug of search engines. It is not duplicate content usually, but rather duplicate URLs for the same physical content.
 
 Suppose you have a page at http://seo.yu-hu.com. Just one physcial page. This one page can be reached in four different ways.
 
That is a simple case  for a site that uses physical files, not pages generated from a database.
A site that is run by a content management system however, may generate the same exact content in dozens of ways, from different URLs from the "products" or "catalog" or "archives" sections. It is still the same physical content that comes from the Database.  
 
Google  and other search engines decide that the additional pages are "duplicate content." - They really are.  
It is not clear how this may penalize your site or if it penalizes it.
 
Google and Yahoo! now let you tell them how to index the page. You do it by putting a "Canonical" attribute  in the head section of the page in a dummy link tag, link this.

<link rel="canonical" href="http://www.example.com/product.php?item=swedish-fish" />
 
The result should be that all the pagerank and other goodies will be given to the version of the page specified. See here for more details.
 
 

Saturday, February 14, 2009

How Google treats Duplicate Content

Duplicate content is a headache for Search Engine positioning. Duplicate content is created in these situations:
  • Syndicated news items that appear in many Web sites
  • Sites that legitimately archive news items
  • Blogs that legitimately quote part of an article, or provide an entire article for reference.and link to the original under "fair use" doctrine.
  • "Classic" articles that are copied at numerous sites.
  • Classic literature and poetry
  • Plagiarism - A takes the content of B without permission and without giving credit to B, the originator, and puts it at their own Web site, or copies it to some Web forum or community Web site.
Very often, online journal articles disappear from the Web after a few years, and journals routinely do not reply to requests for syndication or quoting, so it is legitimate to quote large parts or all of such articles to make sure the record is intact, especially if your web log commentary refers to the article and makes no sense without it.
Plagiarism is a different matter of course.


Different search engines may deal with Duplicate content in different ways. Google uses a patented algorithm for finding duplicate content. It can put all or most duplicate content in "supplementary listings" that will not even be shown unless requested. Others may not even list those pages. The big problem is to determine what page "deserves" to be listed at the top of the SERP (Search Engine Results Page) listing. Most people are going to click on the top listing.


Obviously, the Web site that has the oldest file is probably the originator and should be listed first. But that is frequently NOT what happens. Plagiarism is the sincerest form of flattery. I wrote a rather successful article about an issue. It was promptly copied to a large Web site, and the listing at that Web site pushed the listing of my own page in Google way down in the page. In a different case of plagiarism, material published at our website was copied to a major journal, and to a major news service, neither of which gave any credit for the original and both of which claimed they had copyrighted our material!


If you are only "in business" to influence political opinion, then of course you are willing to sacrifice popularity of your own article in order to spread the word to the largest number of people. But in the long run, you still want your Web site to get more traffic, as that will help your cause the best.


In another case, I looked for an item using a keyword and found a prominently listed page at a closed, keyword protected, Web site. As it turns out, the original article is in the public domain and is freely available at another Web site, but that could only be found by searching the supplementary listings.

Google (if they are listening) should look into this problem, as it reduces the quality of their results, and in the long run, it will reduce the quality of materials on the Web. There is no practical way to prevent copying of materials, but these copies should all credit the original version and link to it. Authors should not be cheated out of credit for their work - that will not promote the creation of quality materials. If you own, or control a Web site or archiving forum, insist that any duplicate material that you copy must link to the original version on the Web. That will not necessarily ensure that the original is listed at the top of SERPs, but it will help. It will also reward the originator by providing the originating site with all important Link

There is one exception - if you are quoting an item as an example of hate propaganda, it seems to you that you are not morally obligated to provide a live Link to the original site and help their website popularity. You can provide the text of the URL without a live link or use a Nofollow attribute

.
Ami Isseroff

Duplicate and Triplicate Google Ad-Sense advertisements

The recession is upon us. That seems to mean that for many topics in many locales, AdSense may display the same advertisement in more than one ad slot on a page. Of course, this reduces Click-Through Rate (CTR- the percentage of visitors to a page who click on advertisements) because nobody will click on the same ad twice, and people who are not interested in finding out how to get a flat stomach might be interested in finding out where to get gourmet foods. Variety of advertisements obviously should increase CTR. Google often puts duplicate ads on a page even when there are different ads (also duplicates) on other similar pages!

The ways that Google seems to use to decide what ads to put on a page according to content are somewhat mysterious. You may have a whole page about astrophysics, but if for some reason there is a single link to a poetry website on that page, they may put an advertisement for poetry there. There algorithm may be a bit primitive.
You would have thunk that if Google AdSense knows about your page content, they also know what advertisements they put there. Since Google gets revenue from the advertisements, they should be interested in maximizing click through rate, right? I have not seen that anyone who obtained an interview with a Google guru asked about this problem.

Until Google acknowledges the plight of their suffering publishers and fixes the problem, you can help fix this condition a bit. Different shaped ad slots will draw different ads (though large horizontal and vertical graphic ads seem to draw the same content. You can also specify that one slat accepts graphics while another is text-only. What a pity that Google has that rigid rule about three ad units per page, whether they are big ad units or little ones, and whether it is a huge page or a little one.

Sunday, February 8, 2009

When did it happen?

Take a look at this news item:
 
Screen resolution 800 x 600 significantly decreased for exploring the internet according to OneStat.com
 
Amsterdam - July 25 - OneStat.com ( www.onestat.com ), the number one provider of real-time intelligence web analytics, today reported that more and more internet users choose for screen resolution 1024 x 768 which is the most popular screen resolution for exploring the internet.
 
The finding has important implications for web site designers because most web sites are designed for a screen resolution of 800 x 600 pixels.
 

The screen resolution 1024 x 768 has reached an all time high and has risen from 54.02 percent in June 2004 to 57.38 percent. Users with monitors set to the most common resolution 800 x 600 for web sites have an approximate 18.23 percent global usage share. A year ago this percentage was 24.66 percent.

Only one detail is missing from this page - the year. When did this happen? We can guess from the last paragraph that the article was published in 2005, but it does not say that.
 
Don't forget to put dates on time-locked materials.
 

Friday, December 19, 2008

IMPORTANT NOTICE - Read Me First - Before you pay for SEO (Search Engine Optimization) Services!!

If you are about to hire a Search Engine Optimization contractor, spend a few moments reading this - it could save you thousands of dollars.
 
Not long ago I wrote a little article about How to choose a contractor for Search Engine Optimization work. Little did I realize the vital importance of this information for the world. 
 
I have since seen quite a few desperate forum pleas by people who have been ripped off by "SEO" contractors. The reason they got ripped off is that they do not know what Search Engine Optimization is. Don't buy on automobile unless you know what automobiles do, and don't pay for "SEO" unless you know what it is supposed to be and what the firm is going to do for you and why. If you understand a bit of what it is about, you will not be ripped off as easily and you will also be able to get the most of any search engine optimization service.
 
The short version:
 
Search Engine Optimization is a bunch of techniques for making your website more visible in unpaid search engine listings. These techniques involve

On Page OptimizationKeywords and coding and content on a page.

Off Page Optimization - This usually refers to Links from other Websites.

Website SEO design - Link structure in your site.

A longer version: SEO Basics and The SEO Mini-Book. You should at least understand the basics before hiring an SEO contractor.

 
What  Search Engine Optimization is NOT:
 
Paid advertisement in search engine results pages or other Websites - Not SEO and you do not need a firm to do that for you.
 
Black Hat SEO - Practices that can get your site banned from search engines. If someone wants to do anything like that or says "this might get your site banned," don't do it.
 
Web page graphic design - Changing the outward appearance of Web pages in itself may have no effect on search engine placement or number of visitors to your site. It might affect conversion rates, which are peripherally related to Search Engine Optimization.
Choosing a Search Engine Optimization Firm
 
Make sure you understand exactly what they intend to do and why - what are you paying for. That's elementary. This should be defined both in terms of operations they will perform and of results they expect to achieve.
 
Operational
Are they going to do work or just tell you how to fix the site?
Are they going to change code on your Web pages?
Are they going to change text on your Web pages?
Are they going to add new Web pages to your site?
Are they going to add new links to your Web pages from oustide?
What key words and phrases will they optimize for your site?
 
Check Before You Buy
Ask to see sites that this SEO firm has done. Do they get good rankings in Google for popular keywords?
 
Do pages have meaningful file names like nice-widgets.htm. Long and meaningless file names and URLs like outasite.com/products/cat/0124sdsfag12893435576fz are a sure sign that there is no search engine optimization here.
 
If they are generating the new site, is it making "dynamic" pages (based on a database and generated on the fly) or static html (actual html files) ?? Static html is better.
 
Does every page link back to the main page using a keyword that people will search for?? ("Home" is not a keyword unless you are selling homes).
 
Check the source code (in any browser, click View-->Page Source). Do images have  "alt" tags with keywords? Is the <TITLE> tag in the header a keyword? Is there a site map for the site?
 
Talk to Webmasters who worked with these people. Did they get more visitors? Check the Web sites in Alexa (http://www.alexa.com/data/details/traffic_details/)  to see if there was really a recent improvement in site rank for traffic.
 
 
Results
Do they promise top spots in Google or other search engines for your pages for different keywords?
Do they promise to increase the  Google PageRank  of your site?  
Do they promise a percentage increase in number of visitors??
 
Beware
Changing names and files - Firms that want to change the domain name or filenames of your site generally do not know what they are doing. Changing your domain will reduce traffic if you have any visitors now. If the name MUST be changed, the old pages and domain should remain and should be redirected to the new ones.
 
Java and Flash - Javascript is a turn off for search engines. You do not want pages with Javascript menus in them or lots of fancy javascript for display graphics or Flash immages. Google is only now learning how to deal with Flash images.
 
Meaningless promises - Top 10 spots for keyword in Google - If the keyword is Sexonomonia, it won't help you because nobody is searching for it. If the keywords is "Maps" they cannot do it for you - too many other good Web sites rank ahead of yours. Be sure you understand what they are promising and why it is good for your site!  
 

Thursday, December 18, 2008

Registering Web pages in search engines - a better way

Here's my year end nondenominational Christmas, Hanukkah, Kwanza and Eid el Fitr wish list from  Google. It's simple:
 
1- Don't penalize me for Google bugs. I really cannot control whether someone links to http://seo.yu-hu.com/ or http://www.seo.yu-hu.com/ or http://seo.yu-hu.com/index.html or http://www.seo.yu-hu.com/index.html. If your spider and software are too dumb to understand they are the same pages, it is not my fault.
 
2- There is really no difference between a page that is in a subdirectory and one that is not. Really! If you are going to penalize pages that are in subdirectories, or index them later, you ought to be telling people that that's how the spider works. If you like flat directories, we will give you flat directories. Just ask!  
 
3- Most important - Give us all a quick and easy way to tell you when I have 1 or more pages or to tell you to index the whole site. There are a dozen good technical ways to solve that problem. The XML Site map is one of the bad ones. If you are going to insist on those maps, then provide a free tool that will crawl the site and submit the URLs in any format you like. But in addition to that, please give me an easy way to tell you that there is one new page at the site. I should not have to make a whole site map for a single page.  For bloggers it's easy. There is an RSS syndication file and that file can be sent to an aggregator. You probably use information from aggregators and blogs like Blogline. But what about sites that do not have an RSS feed? Why isn't there a simple interface for submitting a single URL to a queue? It could be used for new pages or changed pages. That could also take a load off Google software, since widespread use of such interfaces reduces the need for frequent spider crawls through thousands of pages to find just one that is really new or changed. It is incomprehensible why registration of new pages has to be such a hassle when there are simple and foolproof technologies available to solve this trivial problem.
 
Is that too much ask?
 
Anyone out there who sees this cry of despair - post it to Matt Cutt's blog and maybe Santa will answer our prayers.
 
Ami Isseroff
 
 
 

Monday, December 8, 2008

Less Website traffic on Weekends and in Summer

A number of people have complained in forums about the drop in Web site traffic in summer and weekends. It is certainly a fact. There are not a lot of data about this, but there is at least one published set of yearly statistics showing there is a trough every August. I could not find any data about weekends, so I posted some of my own. My own Web sites are more affected than most by summer traffic depression and perhaps by weekend depression as well. Traffic in May is about 1.8 times that in August.
 
Here's the article with the graphs: Weekend and Summer Web traffic depression 
 
If anyone has any suggestions about what ought to be done about it, or ideas about long term trends, this is a good place to discuss that.