Tuesday, January 5, 2010

Cloud Computing: Don't junk your PC just yet

Self-styled financial mayvens and fools, motley and otherwise, are insisting that the "wave of the future" is cloud computing. The bulky and expensive PC will be replaced by a low powered and simple computer that runs a "thin client" software program connected to an Internet Software Services Provider, such as Google. The provider will run the latest versions of spreadsheets, word processors, desk top publishing applications, CAD-CAM programs, Customer Service Management programs, presentation software, database programs and anything else you can imagine, which your business can access for a nominal fee. No more expensive hardware upgrades, no more expensive software upgrades, and every function will be available everywhere on a netbook or equivalent computer. An end to the PC and the beginning of computing nirvana. Well maybe. Or maybe not.
 
If you have nothing to hide, you have nothing to fear... But everyone has something to hide.
 
There are a few flies in the hypothetical ointment of the Computing Cloud. The first is that quite a few firms and private individuals are very content to run hand-me down versions of operating systems and software that are just as good for their purposes as the new updated ones, and do not cost anything at all to upgrade. And though we won't talk about it (not "nice") some people actually steal software. None of those people or firms are going to be interested in paying a "nominal fee" each year for the version of whatever it is that includes the latest bugs, when they have a perfectly fine working version of a Word processor, spreadheet, presentation manager and database, which they know how to use and which is compatible with the files they have already created.
 
Moreover, not all updates can be done smoothly or without changing hardware. Changeover from 32 bit to 64 bit to 128 bit computers and operating systems will require new hardware and new locally installed software. Each major technological change is still going to require junking the old computer, and these will continue to happen every few years to provide better displays better mice or no mouse at all (which would be the best sort of mouse!) bigger and better disks working on new principles, and numerous other innovations. Remember when a computer display took up an an entire desk, in the "bad old days" of five years ago? Or when a 100 MB hard disk was really huge? I remember when 10 MB was a lot of hard disk too. These innovations will all require new hardware no matter what you do. The hardware industry won't stand still, and people will still need some or all of these innovations in the thinnest of hypothetical thin clients.  
 
Another problem is dependence. Everyone who remembers the bad old days of the big central computer also remembers the announcements that flashed across the screen as you were frantically trying to meet a deadline, "Acme central computer will be closing down for two hours in five minutes for maintenance. Please save your work and exit all active programs." The PC freed everyone from that.  Supposedly that won't happen in cloud computing, or will it? And nobody guarantees that the cloud computing provider is always going to provide the latest hardware or the full amount of computing capacity that you need. Everythng has limits and everything cost money, and it is always cheaper to provide less rather than more, and to maintain a capacity that is adequate for most of the day and most of the time, but not necessarily enough for peak use.
 
However, the biggest problem of cloud computing, the show stopper that will make it a no-go for most firms and individuals, is data security. The enthusiasts of cloud computing, asked about security, will spout reams of gibberish about https and unique private keys and prime numbers and 128 bit encryption. Technojargon to awe the uninitiated. The truth is more prosaic. There is no way to protect data against a determined attacker.
 
As a cloud computing executive said, "If you have nothing to hide, you have nothing to fear." But everyone has something to hide. Every firm has their latest designs in their database and word processors. They have all the contact information of their customers in their CSM database. They have the amounts that they bid on various sealed bid contracts. They have the studies that show their own weaknesses vis-a-vis their competitors. All of this is quite interesting information for competitors and industrial spies. All the fancy security protocols and encryption schemes in the world depend on you being you and not someone else, and they are all vulnerable to identity theft. Identity theft is a multi-billion dollar industry and it is growing, despite the best efforts of ingenious firms to foil it. Once the thief has your information, and can make the system think he or she is you, they can access anything you can access. A disgruntled former employee of your firm, a disgrunlted former employee of the cloud computing service, a dishonest employee or a determined industrial spy can and will get the passwords and counter passwords by phishing schemes, by stealing Wi-Fi signals or by hacking databases. There is no system so foolproof that it cannot be fooled. It happens all the time. And what happens to your data privacy if the internal revenue service (or your wife's divorce lawyer) forces the cloud computing service provider to produce the records?
 
Do you really have nothing to hide? Are you sure? I thought so. Don't throw away your PC.
 
Ami Isseroff
  
 

Friday, August 21, 2009

The end of free Internet news?

Rupert Murdoch and others have decided it is time to end free news content  on the Web. One of the reasons that someone like Murdoch could make such a proposition is that he admittedly knows nothing about the Internet or how it works. I think there is no way at all to really end free content and that paid content will never be able to compete with free. Consider firstly all the legitimate primary sources of internet news that are always going to be free: Government Web sites, government broadcasters and NGOs. Governments and NGOs want you to see their content. A large part of international news consists of refurbished government press agency announcements or NGO press releases. "The prominent NGO Birdwatchers International has relased a new study showing..." and the rest of the item just quotes the information in the press release.
 
Now consider also the question of copyright and "fair use." A blogger subscribes to a paid service and copies the main content of an article to their Web log. As long as they comment on the article, and they are not a for-profit organization, it is "fair use for educational purposes." Attempts to stop them will be stymied because they will be labelled attempts to stifle freedom of the press.
 
The claim of publishers that they produce "quality content" that people will want to pay for is also highly dubious. During the Iraq war and the Second Lebanon war and other such events, the press often published government or terrorist propaganda indiscriminately. A CNN report described how dramatic footage of amulances rushing to the rescue was generously faked by Hezbollah for the benefit of the press. A Reuters photo of smoke over Beirut was shown by a blogger to a fake, and bloggers showed many other instances in which the commercial press was fooled by biased stringers or interested parties into passing off fabrications as fact - the French footage of the alleged killing of Muhamad al Dura was one such instance. Consider also stories like Sy Hersh's allegations of an imminent US attack on Iran that never materialized. These stories appeared over and over, though they had no basis in fact. If you want to lie to me for free that's fine, but I won't pay for it. The same was true for Judith Miller's NYT stories about WMD in Iraq.
 
There are so many ways for good free content to get to the Web and be available to all, that it is really doubtful that many people will want to pay for it, especially considering the poor quality of a lot of commercial journalism.
 
Ami Isseroff

Friday, May 22, 2009

Decline of Dmoz: Schadenfreude and sadness

As a long time frustrated user, submitter and ex-editor of the Open Directory (AKA Dmoz) I had feelings of Schadenfreude mixed with sadness when I learned of its decline. It has lost a lot of its viewing public, mostly because search engines do the search job better, but it has also stopped accumulating new listings. Triplicate listing of garbage pages, editors that tyrannize people with other political viewpoints and confine their directories to polemical articles or to their friends' web sites, arbitrary and capricious editing rules used to keep out sites that editors don't like, all detract from the quality of Dmoz. Lack of quality ratings and quality criteria are also a problem.
 
There are a few articles on the decline of dmoz around the Web, and they have attracted quite a lot of comment. Dmoz editors keep writing to say how great they are, without any understanding of what the statistics are telling them. Frustrated users are venting, but dmoz editors are never going to take them seriously, and that's a big part of the problem - contempt for users, arrogance, the elitism of a closed group. But there are a lot of good editors at dmoz and it is worth saving from itself.
 
Some enterprising people made a dmozsucks.org Web site. God bless 'em, but directories like dmoz serve an important function if they are run right, because they can provide information about quality of Web pages to search engines. My detailed thoughts about this are at: The Decline of Dmoz.
 
Ami Isseroff

Sorting media garbage from media information - with special application to the Web and Internet

When I wrote the article: The Decline of Dmoz. it got me thinking about how to give Web page and other media raters objective criteria for deciding if an article or other media item is useful or good or if it is flotsam to be ignored. If there was a directory for "everything" how would you keep out the flood of garbage in internet, printed matter, video and TV, and how could you spot and highlight the really superb new articles or books that should be highlighted and emphasized. It's not as easy as you think. Leonard and Virginia Woolf had a publishing business, and one of the surprising things that they found was that in the long run, the books and poems and articles that were least popular in their initial publication often became best sellers. Indeed, their tiny, romantic, hopeless venture, Hogarth Press, that operated from a hand press, produced some of the greatest classics of the twentieth century. But these great artists sold pitifully small numbers of books when their works first appeared.
 
A scale of quality would be useful for consumers as well, since it would give them a better idea of how much reliance to place in a Web page, article or newscast. For that, we would have to eliminate some of the most obviously useless categories I will mention below, and provide more details of how to judge the less bad material.
 
This scale is going to need a lot of work, but here is a first go, from the bottom (or near it) to the top.
 
Web sites that should not be indexed at all:
 
Web sites that have taken over domain names and use them for porn, gambling or other exploitation.
 
Parked Domains
 
Gimmick sites that are just search engines or advertising
 
Plagiarized material - material that is taken verbatim from another Web site without a link to the original, and often without specifying author or credit and posted  to another Web site, Web log or forum. I have been, and am, the victim of this sort of thing and I am not the only one. The people who do it invariably have an excuse
 
Racist and hate sites, videos etc.  - I think Google's policy is wrong. Web sites like Stormfront, Jew Watch and ihr have no place on the Internet. Spreading disinformation and hate is not doing anyone a service. Inciting to genocide is a crime under the International genocide convention - really it is.
 
Spam and confidence schemes should be banned from the mails and the Internet.
 
Search engines index a surprising quantity of such sites.
 
Lowest Quality Materials
 
Anonymous emails and Web logs or sites that post and re-post copies of anonymous material that never had an author who would acknowledge them. They are almost invariably hoaxes.
 
"news" reports based on anonymous sources and without confirmation - whether they are on the Web or in other media, are of the same approximate quality as anonymous email hoaxes about the latest email
 
Opinion pieces or news items that rely on attribution from non-authoritative sources to establish facts, such as "A guy I met told me that Google no longer uses Pagerank for anything and it is not important." The author didn't say it, and it is probably not true. They hid behind the non-authoritative source to intentionally perpetrate a falsehood. This is done all the time by supposedly serious journals
 
Videos and similar material that are so poorly produced that you cannot hear what people are saying. It is beyond me why people post such things to YouTube.
 
Materials that are just copies of articles published elsewhere, properly attributed. These have some utility especially if the original may be obscure or removed from the Web by the publisher.Pn the Web, it is generally considered legitimate to post whole articles provided you give due credit to the original, and arent just duplicating someone else's Web site to steal their income. Usualy though, it is best to go to the original source and to quote only parts of it, if you can be sure the source will still be there in five years. On the Web, you cannot be too sure.
 
Conspiracy theories that are not verified from other sources. There are whole Web sites devoted to the most fantastic ideas, usually based on total disinformation and often involving race hate and paranoia. The FBI and the Mossad did not cause the 9-11 attacks or the attack in Mumbai. The Federal Reserve system is not a plot to steal your money and give it to rich bankers.
 
 
Materials to be treated with due caution
 
Claims made by commercial sources who are selling a product
 
Sources with an obvious political bias.
 
Materials that use adjectives or hype to describe products or political issues. If words like "right wing" or "left wing" or "progressive" appear too often in an article or report, you have to ask yourself if this person is telling you facts or trying to convince you of their opinion.
 
Differential treatment of subjects - For example, a publication that will regularly use adjectives like "right wing" or "extremist" to describe politicians on one side of a conflict, but refrains from using any adjectives to describe leaders of the other sie.
 
Materials that are unsourced.
 
Assertions from publications or authors who have a poor track record for accuracy. Certain people for example, regularly predict that Iran will explode a nuclear weapon in a few months, or that Israel or the US will attack Iran, but it never happens. If they were ignored, they could not make a living by spreading disinformation in that way.
 
An article or publication that omits important facts that you know to be true is probably trying to create bias.
 
An article or publication that intentionally distorts a quote or lies about a fact, should not be trusted about other facts and assertions.
 
An article or book that has more than a few ellipses ("...") in quotes, is probably distoring the meaning of the quotes. This is a favorite technque of certain politicians, and is useful for dishonest commercial purposes as well.
 
Information you can rely on
 
Source is generally known to be correct
 
Information is confirmed by other reports
 
The report is plausible based on scientific evidence and common sense.
 
Source has no reason to lie
 
There's a lot less of that around than you might think. Remember the Iraq WMD that weren't? The Israeli bio-weapon hoax that was reported in numerous respected journals?

Monday, March 16, 2009

Google Keyword Search Frequency mystery, or "Is Sex going out of style?"

According to  Google's  AdWords tool, sex may be going out of style. That is,  Google's data show supposedly that over a 12 month period, on average there were 124,000,000  (One Hundred and Twenty Four Million) searches for keyword Sex each month, whereas in the month of February there were only 90,500 (Ninety thousand Five Hundred) people searching for sex. No mistake about the number of 0s anywhere either. Double checked.   All the rest of the people looking for sex  must have found it. But sex was not the only keyword affected. Every keyword I checked except Facebook had a lowever search frequency in February than over the last 12 months on average. The size of the drop was not consistent however. Some words dropped much more than others Is search going out of style or there just something wrong with Google's reporting or what?
 

What happened to Jewwatch.com?

A political-social search engine optimization issue developed around the hate web site jewwatch.com. For many years, searches in  Google for the  keyword  Jew returned this odious site at the top of the  listings or among the first ten. Jewwatch.com features standard Anti-Semitism fare including ZOG, the Zionist Occupied Goverment and the forged Protocols of the Elders of Zion. Attempts to get  Google to ban the site failed in the past, but now Jewwatch is gone from the top listings for keyword Jew. The question is:

Sunday, February 15, 2009

Solution for Duplicate page listings - 'canonical' attribute in head tag

One sort of duplicate content happens because of plagiarism, copying to forums, copying of articles to blogs (see  How Google Treats Duplicate Content). In those cases, there are really several physical instances of a page for various reasons. 
 
But there is another sort of "duplicate content" that is often really just an artifact of how the Web works and in part a bug of search engines. It is not duplicate content usually, but rather duplicate URLs for the same physical content.
 
 Suppose you have a page at http://seo.yu-hu.com. Just one physcial page. This one page can be reached in four different ways.
 
That is a simple case  for a site that uses physical files, not pages generated from a database.
A site that is run by a content management system however, may generate the same exact content in dozens of ways, from different URLs from the "products" or "catalog" or "archives" sections. It is still the same physical content that comes from the Database.  
 
Google  and other search engines decide that the additional pages are "duplicate content." - They really are.  
It is not clear how this may penalize your site or if it penalizes it.
 
Google and Yahoo! now let you tell them how to index the page. You do it by putting a "Canonical" attribute  in the head section of the page in a dummy link tag, link this.

<link rel="canonical" href="http://www.example.com/product.php?item=swedish-fish" />
 
The result should be that all the pagerank and other goodies will be given to the version of the page specified. See here for more details.
 
 

Saturday, February 14, 2009

How Google treats Duplicate Content

Duplicate content is a headache for Search Engine positioning. Duplicate content is created in these situations:
  • Syndicated news items that appear in many Web sites
  • Sites that legitimately archive news items
  • Blogs that legitimately quote part of an article, or provide an entire article for reference.and link to the original under "fair use" doctrine.
  • "Classic" articles that are copied at numerous sites.
  • Classic literature and poetry
  • Plagiarism - A takes the content of B without permission and without giving credit to B, the originator, and puts it at their own Web site, or copies it to some Web forum or community Web site.
Very often, online journal articles disappear from the Web after a few years, and journals routinely do not reply to requests for syndication or quoting, so it is legitimate to quote large parts or all of such articles to make sure the record is intact, especially if your web log commentary refers to the article and makes no sense without it.
Plagiarism is a different matter of course.


Different search engines may deal with Duplicate content in different ways. Google uses a patented algorithm for finding duplicate content. It can put all or most duplicate content in "supplementary listings" that will not even be shown unless requested. Others may not even list those pages. The big problem is to determine what page "deserves" to be listed at the top of the SERP (Search Engine Results Page) listing. Most people are going to click on the top listing.


Obviously, the Web site that has the oldest file is probably the originator and should be listed first. But that is frequently NOT what happens. Plagiarism is the sincerest form of flattery. I wrote a rather successful article about an issue. It was promptly copied to a large Web site, and the listing at that Web site pushed the listing of my own page in Google way down in the page. In a different case of plagiarism, material published at our website was copied to a major journal, and to a major news service, neither of which gave any credit for the original and both of which claimed they had copyrighted our material!


If you are only "in business" to influence political opinion, then of course you are willing to sacrifice popularity of your own article in order to spread the word to the largest number of people. But in the long run, you still want your Web site to get more traffic, as that will help your cause the best.


In another case, I looked for an item using a keyword and found a prominently listed page at a closed, keyword protected, Web site. As it turns out, the original article is in the public domain and is freely available at another Web site, but that could only be found by searching the supplementary listings.

Google (if they are listening) should look into this problem, as it reduces the quality of their results, and in the long run, it will reduce the quality of materials on the Web. There is no practical way to prevent copying of materials, but these copies should all credit the original version and link to it. Authors should not be cheated out of credit for their work - that will not promote the creation of quality materials. If you own, or control a Web site or archiving forum, insist that any duplicate material that you copy must link to the original version on the Web. That will not necessarily ensure that the original is listed at the top of SERPs, but it will help. It will also reward the originator by providing the originating site with all important Link

There is one exception - if you are quoting an item as an example of hate propaganda, it seems to you that you are not morally obligated to provide a live Link to the original site and help their website popularity. You can provide the text of the URL without a live link or use a Nofollow attribute

.
Ami Isseroff

Duplicate and Triplicate Google Ad-Sense advertisements

The recession is upon us. That seems to mean that for many topics in many locales, AdSense may display the same advertisement in more than one ad slot on a page. Of course, this reduces Click-Through Rate (CTR- the percentage of visitors to a page who click on advertisements) because nobody will click on the same ad twice, and people who are not interested in finding out how to get a flat stomach might be interested in finding out where to get gourmet foods. Variety of advertisements obviously should increase CTR. Google often puts duplicate ads on a page even when there are different ads (also duplicates) on other similar pages!

The ways that Google seems to use to decide what ads to put on a page according to content are somewhat mysterious. You may have a whole page about astrophysics, but if for some reason there is a single link to a poetry website on that page, they may put an advertisement for poetry there. There algorithm may be a bit primitive.
You would have thunk that if Google AdSense knows about your page content, they also know what advertisements they put there. Since Google gets revenue from the advertisements, they should be interested in maximizing click through rate, right? I have not seen that anyone who obtained an interview with a Google guru asked about this problem.

Until Google acknowledges the plight of their suffering publishers and fixes the problem, you can help fix this condition a bit. Different shaped ad slots will draw different ads (though large horizontal and vertical graphic ads seem to draw the same content. You can also specify that one slat accepts graphics while another is text-only. What a pity that Google has that rigid rule about three ad units per page, whether they are big ad units or little ones, and whether it is a huge page or a little one.

Sunday, February 8, 2009

When did it happen?

Take a look at this news item:
 
Screen resolution 800 x 600 significantly decreased for exploring the internet according to OneStat.com
 
Amsterdam - July 25 - OneStat.com ( www.onestat.com ), the number one provider of real-time intelligence web analytics, today reported that more and more internet users choose for screen resolution 1024 x 768 which is the most popular screen resolution for exploring the internet.
 
The finding has important implications for web site designers because most web sites are designed for a screen resolution of 800 x 600 pixels.
 

The screen resolution 1024 x 768 has reached an all time high and has risen from 54.02 percent in June 2004 to 57.38 percent. Users with monitors set to the most common resolution 800 x 600 for web sites have an approximate 18.23 percent global usage share. A year ago this percentage was 24.66 percent.

Only one detail is missing from this page - the year. When did this happen? We can guess from the last paragraph that the article was published in 2005, but it does not say that.
 
Don't forget to put dates on time-locked materials.
 

Friday, December 19, 2008

IMPORTANT NOTICE - Read Me First - Before you pay for SEO (Search Engine Optimization) Services!!

If you are about to hire a Search Engine Optimization contractor, spend a few moments reading this - it could save you thousands of dollars.
 
Not long ago I wrote a little article about How to choose a contractor for Search Engine Optimization work. Little did I realize the vital importance of this information for the world. 
 
I have since seen quite a few desperate forum pleas by people who have been ripped off by "SEO" contractors. The reason they got ripped off is that they do not know what Search Engine Optimization is. Don't buy on automobile unless you know what automobiles do, and don't pay for "SEO" unless you know what it is supposed to be and what the firm is going to do for you and why. If you understand a bit of what it is about, you will not be ripped off as easily and you will also be able to get the most of any search engine optimization service.
 
The short version:
 
Search Engine Optimization is a bunch of techniques for making your website more visible in unpaid search engine listings. These techniques involve

On Page OptimizationKeywords and coding and content on a page.

Off Page Optimization - This usually refers to Links from other Websites.

Website SEO design - Link structure in your site.

A longer version: SEO Basics and The SEO Mini-Book. You should at least understand the basics before hiring an SEO contractor.

 
What  Search Engine Optimization is NOT:
 
Paid advertisement in search engine results pages or other Websites - Not SEO and you do not need a firm to do that for you.
 
Black Hat SEO - Practices that can get your site banned from search engines. If someone wants to do anything like that or says "this might get your site banned," don't do it.
 
Web page graphic design - Changing the outward appearance of Web pages in itself may have no effect on search engine placement or number of visitors to your site. It might affect conversion rates, which are peripherally related to Search Engine Optimization.
Choosing a Search Engine Optimization Firm
 
Make sure you understand exactly what they intend to do and why - what are you paying for. That's elementary. This should be defined both in terms of operations they will perform and of results they expect to achieve.
 
Operational
Are they going to do work or just tell you how to fix the site?
Are they going to change code on your Web pages?
Are they going to change text on your Web pages?
Are they going to add new Web pages to your site?
Are they going to add new links to your Web pages from oustide?
What key words and phrases will they optimize for your site?
 
Check Before You Buy
Ask to see sites that this SEO firm has done. Do they get good rankings in Google for popular keywords?
 
Do pages have meaningful file names like nice-widgets.htm. Long and meaningless file names and URLs like outasite.com/products/cat/0124sdsfag12893435576fz are a sure sign that there is no search engine optimization here.
 
If they are generating the new site, is it making "dynamic" pages (based on a database and generated on the fly) or static html (actual html files) ?? Static html is better.
 
Does every page link back to the main page using a keyword that people will search for?? ("Home" is not a keyword unless you are selling homes).
 
Check the source code (in any browser, click View-->Page Source). Do images have  "alt" tags with keywords? Is the <TITLE> tag in the header a keyword? Is there a site map for the site?
 
Talk to Webmasters who worked with these people. Did they get more visitors? Check the Web sites in Alexa (http://www.alexa.com/data/details/traffic_details/)  to see if there was really a recent improvement in site rank for traffic.
 
 
Results
Do they promise top spots in Google or other search engines for your pages for different keywords?
Do they promise to increase the  Google PageRank  of your site?  
Do they promise a percentage increase in number of visitors??
 
Beware
Changing names and files - Firms that want to change the domain name or filenames of your site generally do not know what they are doing. Changing your domain will reduce traffic if you have any visitors now. If the name MUST be changed, the old pages and domain should remain and should be redirected to the new ones.
 
Java and Flash - Javascript is a turn off for search engines. You do not want pages with Javascript menus in them or lots of fancy javascript for display graphics or Flash immages. Google is only now learning how to deal with Flash images.
 
Meaningless promises - Top 10 spots for keyword in Google - If the keyword is Sexonomonia, it won't help you because nobody is searching for it. If the keywords is "Maps" they cannot do it for you - too many other good Web sites rank ahead of yours. Be sure you understand what they are promising and why it is good for your site!  
 

Thursday, December 18, 2008

Registering Web pages in search engines - a better way

Here's my year end nondenominational Christmas, Hanukkah, Kwanza and Eid el Fitr wish list from  Google. It's simple:
 
1- Don't penalize me for Google bugs. I really cannot control whether someone links to http://seo.yu-hu.com/ or http://www.seo.yu-hu.com/ or http://seo.yu-hu.com/index.html or http://www.seo.yu-hu.com/index.html. If your spider and software are too dumb to understand they are the same pages, it is not my fault.
 
2- There is really no difference between a page that is in a subdirectory and one that is not. Really! If you are going to penalize pages that are in subdirectories, or index them later, you ought to be telling people that that's how the spider works. If you like flat directories, we will give you flat directories. Just ask!  
 
3- Most important - Give us all a quick and easy way to tell you when I have 1 or more pages or to tell you to index the whole site. There are a dozen good technical ways to solve that problem. The XML Site map is one of the bad ones. If you are going to insist on those maps, then provide a free tool that will crawl the site and submit the URLs in any format you like. But in addition to that, please give me an easy way to tell you that there is one new page at the site. I should not have to make a whole site map for a single page.  For bloggers it's easy. There is an RSS syndication file and that file can be sent to an aggregator. You probably use information from aggregators and blogs like Blogline. But what about sites that do not have an RSS feed? Why isn't there a simple interface for submitting a single URL to a queue? It could be used for new pages or changed pages. That could also take a load off Google software, since widespread use of such interfaces reduces the need for frequent spider crawls through thousands of pages to find just one that is really new or changed. It is incomprehensible why registration of new pages has to be such a hassle when there are simple and foolproof technologies available to solve this trivial problem.
 
Is that too much ask?
 
Anyone out there who sees this cry of despair - post it to Matt Cutt's blog and maybe Santa will answer our prayers.
 
Ami Isseroff
 
 
 

Monday, December 8, 2008

Less Website traffic on Weekends and in Summer

A number of people have complained in forums about the drop in Web site traffic in summer and weekends. It is certainly a fact. There are not a lot of data about this, but there is at least one published set of yearly statistics showing there is a trough every August. I could not find any data about weekends, so I posted some of my own. My own Web sites are more affected than most by summer traffic depression and perhaps by weekend depression as well. Traffic in May is about 1.8 times that in August.
 
Here's the article with the graphs: Weekend and Summer Web traffic depression 
 
If anyone has any suggestions about what ought to be done about it, or ideas about long term trends, this is a good place to discuss that.
 
 
 

Sunday, December 7, 2008

Add sense: Google Adsense in a recession

If you have Google  AdSense ads in  your Website then depending on the category of your site, you may have noticed that CTR (click through rate) dropped for a while - quite a bit, and then seems to have recovered. Less people may be clicking ads in a recession - they aren't buying. And Google may have less ads in stock for that category in each location, so they were putting up multiple copies of the same advertisement or public service ads. Of course, multiple copies of the same ad are not as effective as different advertisements that attract different clientele.
 
Google seems to have fixed that a bit by being a bit more broad minded about what ads they will match to what keywords, or in some other way, because CTR has improved slightly. At the same time, they are trying to make up for low CTR in Webpage ads apparently by generating more ads on search pages. Search pages seem to have a much higher CTR than website pages. Google may be trying to improve their numbers for the next quarter earnings report. In any case, Google is number 1 with about 80% share of advertising revenue. (See Google's Ads Per Keyword Output In High Gear)

Troubleshooting Search Engine Optimization - back to basics

While many people get carried away by Web 2.0 and twitter and arcane marketing strategies, they often neglect the basics. Then they don't understand why their Web site/page is not listed. Optimization "experts" ask in forums why their customer's page isn't doing well. Usually it is because they neglected simple stuff or simply ignored all the SEO Basics.   If your Web page/site is not doing well in Google and not attracting vistors, check the simple minded basic things first:
 
1. Is the page or site listed in Google at all? Doh.
 
2. Links -  Is the page/site  linked from a main page in your Web site, one that has high Google PageRank ?? Is it linked from other Web sites and Directories?
 
3. Does it use your  Keywords in all the important places?  Remember, the search engine spider is a dumb machine. If your page is about Widgets, you have to tell the spider it is about widgets in language that it understands. If you are writing about Widgets in a blog, does the blog article title say "Widgets?" Or is it a fancy title that nobody can associate with your real topic, like "To be or not to be?"
 
4. If you can include some keywords in the filename of the page that can help.
 
5. Did you do the drudgework of filling out the Title tag and the Description Meta tag and the Keywords meta tag in the Head Section ? Is there a clear <H1> title in the page itself (only one) ? Are there a few pages (at least 1) linked to that page from outside your Web site, using the keywords of that page as the anchor text? If you linked to the page with the Anchor Text  "More" or "Read about about it Here" the Google spider is very linkely to classify the page under "more" or "read" - you need to use the title text. The search engine will also conclude that your page is about "more" - no kidding, I saw this happen.  If the Title tag  (<TITLE>This is the Title</TITLE>  in the  Head Section  of the page is blank or says "New Page," you can't expect the search engines to do very well at figuring out what it was about.
 
6. Does the Body  TEXT of your page use the Keyword a lot? If your page is supposed to be about widgets, but all the text is about Brittney Spears, search engines can't know that it is about widgets.
 
7. Did you use alt tags  in your pictures of widgets or whatever your site is about to tell search engines and people using text browser (if there are any) and impaired people that these are all pictures of widgets?
 
8. Did you use Title attributes in the links so that search engines can "see" that these are links to articles about widgets. It is not your fault if someone else named their article about widgets "What is do be done?" or "A very excellent solution."  But if you have too many links with that sort of text, the search engines will think your page is about Shakespeare of the New Testament.
 
9. Did you check the code to make sure that the tags for the Body and Head Section are correct? If there is no closing </head> tag or no opening <body> tag the browser might show the page just fine, but the search engines can't tell that what is there is text you want people to read. Likewise, there can be other junk that the search engine spiders cannot see.
 
10. Is the page mostly text that search engines like, or is it filled with Javascript and Flash jibberish and CSS style directives that should be in a separate file? Is the code clean, or does it have a lot of junk in it that is produced by Wysiwyg editors like <Span style = "style1"></span> and
<FONT size = "2" Face = "ARIAL" Color = "Red"> </FONT> <FONT size = "3" Face = "ARIAL" Color = "Blue"></FONT></FONT>. Wysiwyg editors will fill whole Websites with that sort of meaningless flotsam if you don't edit it out. Search engines see that and lower your score.
 
11. Did you link to the main page from every page of your Web site, using an Absolute Path and proper keywords in the Anchor Text. If your home page is about Widgets, the anchor text must say something about widgets, not "Home." if you link to the main page with the anchor text "home" then you may expect to find that page when you search for "Home" in Google.
 
Nine times out of ten, Web page search engine visibility can be improved greatly just by checking and fixing those  things. And yes, you can go through every page that I or anyone else makes and find ways to improve that page, especially if it has been around for a few years and technology changed.
 
Beyond that, there are always unknown factors and gotchas, the Job factor in SEO. Remember Job from the Bible? You can do everything right and still the Gods do not smile. Remember that though not every page can be linked from the highest ranking page on the site (usually the main page), pages in the lower level sections may take much longer to become visible in search engines. Using nice orderly separate subdirectories seems to hurt even worse, though it should not. Sometimes these subsections cannot be avoided. Make up for it by linking to sub pages from articles and from auxiliary Web logs and site.
 
If all else fails and the pages won't get listed no matter what you do, submit a Site map to Google Webmaster central. It won't guarantee high placement, or even registration, but at least Google will consider the page.  
 
Ami Isseroff
 
 
 
 

Wednesday, December 3, 2008

Showstopper Google bug?

All search engines are based on the concept of Authority of Web pages and Websites. The pages and sites that are deemed to have the highest authority are retried in SERPs (Search Engine Result Pages) at the very top. Google based its success on having the best measure of Website authority - the Google PageRank algorithm. The rationale for this intellectual and technical feat and the mechanism are descibed here: The PageRank Citation Ranking: Bringing Order to the Web.
 
Let's see how good it is. The best authority on the Web for the what it says in the Bible is a copy of the bible, no? Anyone who quotes from the Bible might make a mistake, but the Bible is infallible about the Bible, I would think.
 
Here is a quote from the King James Bible:
 
If then God so clothe the grass, which is to day in the field, and to morrow is
cast into the oven; how much more will he clothe you, O ye of little faith?
(Luke 12:28)
 
The phrase "O ye of little faith" appears in the book of Matthew as well. I searched for this phrase, in quotes in Google. The first references that were quotes from the Bible appeared in the eighth page of results! Items like Time magazine articles, song lyrics and an article in one of my Web sites accounted for the first 70 or so results. Google claims it has about 36,000 results for this phrase.
 
"To be, or not to be" is the overly famous quote from Shakespeare's Hamlet. When this is searched in Google, a lone result appears in the third place in my part of the world, behind two Wikipedia articles. It is not from a Web site with the whole play, just a fragment with the soliloquy. The next result that is really from Shakespeare's play appears around position 75 again. Is there a 70 penalty for having the original text (like the 30 Penalty)? Google has about 2.5 million pages with this phrase, or so it claims.
 
Part of the problem is that we used phrases that are extremely popular. I looked up "which is to day in the field"  in Google and indeed, the very first page retrieved was from the Bible. But it was the only one actually from the Bible on that page! There were about 25,000 such pages in Google.
 
I also tried a different phrase from the same Shakespeare soliloquy, "But that the dread of something after death" - not all that famous. Not a single one of the first 10 results was a link to the original Shakespeare text. The text of the "Tragedie of Hamlet" was first listed as result number 26! There were 23,700 results for this phrase.
 
If we cannot rely on "authority" to get search engines to deliver the authentic and authoritative origins of quotes at the top of search results, then it doesn't seem to be worth that much.
 
I had better luck with this line "somewhere i have never travelled,gladly" - it is the title of a poem by e.e. cummings, who is apparently not quite as famous as the Bible or Shakespeare, and therefore he is allowed to be more of an authority on his own work than William Shakespeare or the King James Bible. For this quote too, there were about 9,000 pages claimed by Google. The entire first page of results and more were filled with links to the poem itself.
 
For "How do I love thee? Let me count the ways," the start of Sonnet 43 from Elizabeth Barrett Browning's "Sonnets from the Portuguese" the first three entries retrieved by Google were the actual poem. That's fair enough. More than that would be useless. Someone might be searching for a different page. 
 
  
 
 
 
 
 
 

Phishing - Black Black Hat

Something a bit different this time. Phishing scams and the like are probably the ultimate in  Black Hat SEO and they are a real and costly danger to e-banking and any Web sites that allow financial transactions. (see Danger! Your identity is not secure). Phishing uses a variety of techniques to direct victims from a legitimate Web site to the scam Web site, where they enter their identification information, account numbers, social security, credit card numbers etc. that are then used by crooks to steal their money.  Other schemes download appropriate  Scumware to the victim's computer. This captures their identification details and sends them to the crooks who operate these schemes.
 
Scientific American has a useful article on How to foil Phishing Scams, However, it is only useful up to a point. Most people refuse to become technically educated, and if they do, the scammers will just find new techniques to beat the system.
 
Banking and other institutions have adopted various security measures such as token devices, but these are difficult to manage for various reasons. Identiwall has a promising system based on multiple authentication using cell-phones. It can be applied for Secure ebanking, online stock brokerages and any other type of Web operation, financial or otherwise, that requires confidentiality.
 
The ultimate phishing scheme is one that can steal a legitimate Web site and redirect visitors through Cloaking, Shadow domains and similar techniques. Visitors think they are giving their identification information to the bank or stock broker, but the crook is at the other end. In principle, it is not morally or technically different from Black Hat SEO.
 
 

Monday, December 1, 2008

The nemesis of search engines

The logic of dialectics dictates that everything carries within itself the seeds of its own destruction. The basis of Web search engines is  Website Authority and Web page authority, determined by age and incoming links. That is also the basis for  The killer search engine bug. That's the short version. If you haven't figured it out: Older is not necessary better. An older description (say 1960) of how a computer works and what it can do, is not better than a new one. But the authority algorithm favors older pages. Bigger is not necessarily better either. Apple vs IBM mainframes. But authority algorithms favor the bigger Web sites. See The killer search engine bug for some of the gory details.
 
 

Sunday, November 30, 2008

How to choose and not choose a Search Engine Optimization Consultant

I saw a Website that raised the issue of how to choose a Search Engine Optimization firm. It looked very promising, but in fact, I didn't find it particularly helpful. So I wrote an article of my own. The meat of that article is 11 things you must know before choosing an SEO firm. Why 11? Because it wouldn't fit in 10.  The most important things:
1- Understand when a Search Engine Optimization Firm can't help
2- Understand what things you need to do yourself.
3- Make sure you know exactly what SEO is
4. Make sure you know exactly what this SEO consultant is going to do
 
The rest and some comments on that other site are here Selecting a Search Engine Optimization Firm
 
 

Conversion and Landing Pages

In writing about Landing Pages, I studied several articles that Google found for this search phrase. Remarkably, almost all of them were really about Conversion Rate! A Landing page, for those who do not know, is the page that is meant to be an an entry way into the Web site. It is the "bait" in commercial Web sites. Conversion rate is the measure of how many suckers visitors go from this landing page to the page where they sign up, buy something or do whatever it is you want them to do.  In this antiquated Website model, there are no search engines out there. All your traffic comes from paid advertisements or emailing, and therefore you know where the visitor is going to land and how they got there. So all the advice about landing pages, not surprisingly, is about conversion rate, which, when you think about it, is about how to get people off the landing pages and n to other pages.
 
At the other extreme from the believers in the Landing Page model, are those who believe that "every page is a landing page." Theoretically it might be true in the age of search engines, since there are less "inner" pages - the old model of the hierarchical Web site with a single front page entry point is long dead. But literally it is absurd. Not every page is a landing page. A custom 404 error page is not meant to be a landing page. A "Thank you for filling out the form" page is not a landing page, a print article page is not intended as a landing page.
 
But...
 
But even in "organic SEO" it is impossible and a waste of effort to make every page in the Web site a top draw. Some pages are more equal than others and should get more effort and thought. They may be portals such as an entry to a product list or list of links, or they may be very well written and researched articles or pages with stunnng or interesting graphics. Sometimes a page becomes unexpectedly popular. There is generally a reason that is obvious after the fact (not always).
 
But in analyzing the statistics of any website you can see that there are "landing pages" in the sense that some pages get a disproportionate about of the entry traffic to the Web site, and these are not necessarily the ones you expected to be top pages. In a site with several thousand pages, only about 550 were entry pages in a given period. Among those, the top 20 accounted for about 58% of the traffic! That doesn't mean the site would have 57% of its current traffic if it had only those pages.. The less popular pages support the more popular ones by linking to them, adding to the pagerank of the site etc. But those 550 pages are the ones that are delivering the message to visitors of that site, and the first 20 remain about the same month after month. So you need to check site statistics and exploit the opportunities. If visitors are getting to a page that isn't really what you want them to see, you have to figure out how to exploit that and lead them to a page you do what them to see.
 
How to get people to a landing page is the topic not covered in articles about landing pages. The answer is the same answer as for all search engine optimization, because SEO is optimization of landing pages:
 

On Page OptimizationKeywords and coding and content on a page.

Off Page Optimization - This usually refers to Links from other Websites.

Website SEO design - Link structure in your site, Type of software used to create your Web pages, size of your site...

 
As for improving Conversion Rate, I put most of the important tips I found plus some of my own in the Conversion Rate article, along with a lot of useful links. There are some important recommendations that can help you, some obvious bending of corners, and some very honest advice that says, "These factors can influence search rate. However, they work differently for different products and situations. Therefore, you have to try different ideas and see what works (A/B or multivariate testing).
 
A big problem you always have to deal with is that the things that make for good conversion rate often make for lousy Web pages from the standpoint of search engine optimization and user navigation. The ideal landing page has a fairly ugly and prominent message, like
 
"Buy superwidgets and gain immortality,
 riches and unlimited sex - Guaranteed or your money back"
.
And it has a single button or repeated button or link:
 
Don't Wait!
Click here to Buy Superwidgets NOW!
Hurry while they last!
 
It doesn't have a lot of Long tail  content for search engines to glom on to and it doesn't have navigation links and fancy doodads.
 
I trust you will find a lot of things to think about and use in those two articles about  Conversion Rate, and Landing Page and some ironic laughs, like the SEO "expert" who had to hire someone else to fix their landing page, and the breakthrough conversion recommendation that resulted in a huge increast in conversion rate but not a single new sale. How could that be? Think about it...
 
Meanwhile...
 
Don't Wait!
 
 
Ami isseroff