Showing posts with label problems. Show all posts
Showing posts with label problems. Show all posts

Friday, April 4, 2008

eNets payment is so unfriendly + Vista sucks

Previously, I heard from my friend that the back-end for eNets payment gateway was "upgraded". Not sure how it went, but today is the first time I experienced it from the "front" when trying to pay for some coach tickets.

I am using my Vista laptop with Firefox as the web browser.

1st Attempt
Upon checking out my shopping cart, I was redirected to eNets' Welcome Gateway, which required me to enter my email address (why? no reason? happy? or so they can blast marketing material at me? do they have the right?), and my choice of my payment. Hey! If there is only 1 choice of payment, do I still have to select it from a dropdown list? Why so mafan? Why can't they just detect it at the server side and then present the item to the user?

Nevermind, I submitted my details and selection to the server and was promptly presented with the next page. I tried to enter text into the Name textbox, but it was disabled. I can't enter any data at all!

Seems like eNets do not like Firefox, or because I'm not using the latest and greated Java 6 JRE (where got payment gateway ask you to upgrade your bloody JRE in the middle of a credit card transaction one? this must be the World Class syndrome, you want World Class service, you better be prepared to have world class JRE installed!)

So ok, I cancelled the transaction, no choice. Went out, closed my login session on the coach website. Started IE 7 on Vista.


Attempt Duo,
IE7 started up, I went inside through the same action. Luckily the coach website kept my shopping cart across session! (YAY!). Went to eNets website, duly typed in my credit card details, and then clicked the "Submit" button.

*BOOM*

Vista has halted all my web surfing for no reason. I can't go to yahoo mail, gmail, blogspot, etc.. Nothing! All the pages time out on me. MSN messenger is still running fine. This is not the first time my Vista has done this to me, sometimes all TCP connections get shutdown too. In fact, on my Vista laptop, if I even plug in a LAN cable into the port it will blue screen immediately and die. Yes, this is my Vista experience so far. It sucks. Today, the suckiness has dropped to a new low. It hasn't been this low in ages. Thanks alot, Bill Gates, and take your fishes along when you go. I'm not impressed by Vista lor...


Attempt 3
Really fedup. Rebooted laptop. Logged onto my Windows 2003 Server R2 SP2, 64-bit OS. Started up the IE7 on the server.

Login to the coach website, ah ha! my cart is still intact! no need for retyping all the details and selecting the seats on the coach. YAY!

Checked out the cart, go go go!!! Redirected to eNets.. and...

*engine dies*

The redirection failed. It just hanged there. I can't even get into the page to fill in my email and select the only payment choice from the dropdown list.

*pfffft* grrr!!!


Attempt 4
My Vista laptop is back, while waiting for the startup activities to finish (yes, after you login, Vista takes another 3-5 minutes doing its own thing, the harddisk activity doesn't stop as it starts up the rest of the stuff in the background. I have since disabled autostart for mysql, mssql 2005, etc. Even though the starting up is fast, it just postpones the actual work to after you sign in, bleah...).

Anyway, I digress... I started up my IE7 and went through the usual, familiar motions ALL OVER AGAIN FOR THE FIRST TIME FOR THE LAST TIME, and this time it works. I was able to get my payment processed and confirmed.


*whew* Imagine if this was the opening day for some blockbuster hit, or if I was trying to book a great seat for my family for the F1 race. This whole fiasco would ruin my chances to get the seat that I want, man...

eNets, the service is pathetic, and instead of providing a service to the user by adapting yourself to the user's environment (browser type, browser version, java or the lack of it), you force them to do it your way. Must be IE7, must have JRE, must be patient to use your crap. Next time I encounter this, I will go down to the shop and pay for it. I do not see the increase in service level that corresponds to the increase in your payment processing surcharge.

Extremely horrible experience...

Friday, February 15, 2008

World Class infrastructure for a World Class Event?

No, it wasn't to be so...

This is the headline from Straits Times article

Website booted him out three times
British Airways pilot Benterman takes 10 hours to get tickets for F1's first night race

To make it a double whammy, the permanent resident found, to his horror, that the seats he had reserved were lost when he was booted out of the website.

'It was absolutely frustrating and a disgrace,'' said the exasperated 39-year-old.

'I cannot accept not getting through to the website because it crashed. There is also no customer service number to call.''


Apparently, the website was supposed to be capable of handling 20,000 transactions per hour and the actual traffic apparently was way over what was expected.

Questions:
  1. Did anyone take the last F1 race's figures for a comparison and benchmarking?
  2. Was there a big change in the way the tickets are sold?
  3. Did the system incorporate proper transactions handling, queueing and all that?
  4. Was the system properly load tested before going live?
  5. Was someone even monitoring the system after it went live? Why was no action taken? (cf. MRT down, buses were deployed)
  6. Was there a contingency plan in place? (obviously not)
  7. Could the launch have been scheduled in phases? (online sales first, then outlet sales?)
  8. Were corners being cut in the system hardware so someone could save a few bucks? Or was the organize scammed by vendors who gave 3rd class hardware for 1st class prices? (which is normal)
According to the news on the radio this morning, a hardware upgrade should be sufficient to solve the problem. Shame on the SI (system integrator) and/or hardware vendors who supplied the "solution".

So, we'll see... :)

Monday, February 11, 2008

Can? How much? How fast?

When doing freelance IT projects, some questions from prospective local customers would sometimes be like this:

"Hi, can you do a SQL/web/intranet/(fill in with appropriate IT word) program?"
"How much har?"
"When can finish?"

This is very common and I am always very cautious when dealing with such customers because:
  1. They typically (>90% of the time) do not know what is the effort involved (most likely they learn of this requirement from an in-house IT guru, who might or might not have any experience in IT)
  2. They also do not know what they really want (they are just relaying someone else's words)
  3. They only look at the cheapest quote
Most of the time, I would quote them what I feel is reasonable (of course!) base on a lot of assumptions (the more assumptions and buffer, the more costly it is) as they can't provide me with a reasonable basis to work out a quotation. This is really a case of "you get what you pay for".

For customers whose only concern is cost, I would happily give them a miss as they are the "don't care, don't know, don't bother me unless it is delivered and working" type. These companies are the type that hobbles along on a barely working/functional and often broken IT infrastructure, going from vendor to vendor/supplier whenever the system is down, because they only look at the cost.

Vendor after vendor apply various patches, workaround and upgrades to the original system until it is barely recognizable, and maintainable. Most, if not all, of the time such companies do not have a documentation of the system and the database, resulting in tremendous efforts in tracing through the system and trying to figure out what it is supposed to be doing. And yes, most of the time this work is being performed on production systems too.

Despite the claims by the press, internet and local authorities, the majority of SME owners are rather IT illiterate and clueless (this is base on my limited experience). Most of the time, cost is the only concern. Some of them are happy with a halfway broken system because of the "I know it's broken but I have a staff doing it, fixing it will cost a lot of money lehhh..." way of life. They will devote a staff or two, or even three to perform some of the functions that the system should be performing alone if it was not broken.

Even worse are those who are semi IT-literate, certified as literate after attending a 3 day course in IT conducted by instructors who have barely have any experience in real life IT operations. They are adamant that their "IT Way" is correct and you should not attempt to advise them because they know better.

Let me provide an analogy. Suppose you need to buy shoes, you can either chose to buy a cheap one, or a slightly more expensive but durable one. So, would you rather buy a $30 pair of shoes that spoils every 3 months, or a $150 one that can last you at least a year or more? Some companies would avoid paying the $150 like the plague because the perceived cost is "high". That's sad, limited, and yet very real.

Therefore, in order to secure the job and to fix the problems properly, a lot of communication and persuasion is necessary. The customers must be convinced of the value of the solution, and invest his/her own time into it to help shape the final solution. IT is central to a lot of their operations and can be made to provide more assistance to their business, but yet is given very little priority and investment.

Having said that, there are a lot of moonlighters out there who over promise and under deliver, causing this vicious cycle to continue. I have personally seen some local e-commerce sites with extremely poor exception handling in their purchase and payment code. Yes, they have the usual certificates, logos and seals, but the certification process does not include the testing or validation of the source code itself.

Shop on local websites? Err... maybe not yet... let someone else be the guinea pig for these eBay wannabes :)

Maybe I will get a chance to help fix these borken sites once they get complained :P haha!

Wednesday, January 30, 2008

the mystery of the borken server, SOLVED

Acknowledgements

Thanks to maxsec (from MS's irc channel) and Jules (creator) of MailScanner!


Summary

Problem was two-fold:

  1. I did not notice that Mail::ClamAV and Mail::SpamAssassin packages were not installed properly when running the install script provided in install-Clam-0.92-SA-3.2.4.tar.gz (error information below)
  2. My system had /tmp mounted as "noexec" (is this a default BlueQuartz setting, or did I change this when the system was hardened previously?)


MailScanner diagnosis procedure

  1. After installing MailScanner, run MailScanner --lint, check for any errors that get thrown out.
  2. If there is any issue, run MailScanner -v to see the versions of the installed modules, make sure that they are correct.
  3. If it is not conclusive, run MailScanner --debug or MailScanner --debug --debug-sa (if you have SpamAssassin)
  4. If problem persists, Google it and search the MS Mailing List Archive (it is active).
  5. If there is still nothing conclusive, go to the IRC channel and ask for help.
  6. Subscribe and post the problem in the MS Mailing List too.


Error Information

An error was thrown during installation of Mail::SpamAssassin when I ran the install script in
install-Clam-0.92-SA-3.2.4.tar.gz. (Remind myself to maximize the Putty screen next time).


Setting a soft-link from spam.assassin.prefs.conf into the SpamAssassin
site rules directory.
spam.assassin.prefs.conf is read directly by the SpamAssassin startup
code, so make sure you have a link from the site_rules directory to
this file in your MailScanner/etc directory.
Perl could not find your SpamAssassin installation.
Strange, I just installed it.
You should fix this!

Making backup of pre files to /tmp/backup.pre.3457.tar
tar: *pre: Cannot stat: No such file or directory
tar: Error exit delayed from previous errors
Now go and find your v310.pre and v320.pre files,
echo which may well be in the /etc/mail/spamassassin directory.
You need to save a copy of your old v320.pre file and rename
the v320.pre file to v320.pre.


Moving on

*sigh* :)

Now I have to keep reminding myself to be extra careful when updating this server in the future. Not sure about why the other servers are fine. Maybe the manual installation of SpamAssassin source helped but I didn't do it for this server due to its custom configurations.

In the future updates of MailScanner, I will need to:
  1. Download and unpack the new MS package / installer.
  2. Go into the perl-tar directories and list all the PERL modules.
  3. Open up CPAN (perl -MCPAN -e shell) and compare the version of the installed modules vs those with the MS package / installer.
  4. If the versions are not ok, unpack those files that came with the MS package / installer, manually update them via the usual perl Makefile.PL -> make -> make test -> make install as root.
Ok, that's it for now!

Thanks to the advice and help from the people in MS's IRC channel, and especially to maxsec and Jules!

Monday, January 28, 2008

the mystery of the borken server

Summary

The MailScanner processes on one of my server hangs, and it gets worse as the number of children is increased. Setting a very low number of Children helps, but the problem is not solved.


Background


Server Hardware (Dell)
  • CPU: AMD Dual Core Opteron (2210)
  • RAM: 2GB
  • 2 x 160GB SATA (configured with software RAID 1)
(Key) Server Software


The Problem

The customers (actually it is the customer of my customer) are fairly new, less than a year.

After a recent upgrading, the customers noticed a slowdown in the performance of the email server. Outgoing emails takes a long time to be sent after they hit the "Send" button in the email client. Sometimes it take up to 5 minutes.

So, the parties involved are:
us <-> customers <-> end-customers


Some Context Information

The end-customers are actually located in another country, but the email server is hosted and administered locally.

Network from end-customers to here is routed overseas (which could be contribute to instability at times).

Number of customers is not high, but the network is critical to their international operations.


The Conjecture/Guesses/Hypothesis

  1. Network is unstable or packed, causing upstream traffic to be slow (retrieving emails is fine though). Or their bandwidth is asymmetrical, with upload speed a fraction of the download speed.
  2. Data center network is unstable or does not have peering with customer's network provider, resulting in traffic being routed here indirectly.
  3. Server is under DDoS / spammer attack.
  4. Customer's network has p2p applications running, thereby causing bottlenecks in their internal networks. Or they are hosting web applications in-house, causing their outgoing traffic to be swamped.


Initial Observations

After logging on to the server in the dead of the night (with only some cats and cars passing on the street outside), I noticed that
  • the server load is high, with uptime of >3 (using uptime and top)
  • the email traffic is almost non-existent
  • only 1 user was accessing the server, as evidenced by the paucity of "pop3-login"s in /var/log/maillog
  • MailScanner --lint did not give any errors or warnings
Doesn't seem like the server was under attack (after checking with netstat, lsof), there was spam coming in, at least 1 per 2-3 minutes.

I looked at the MailScanner process and found that it was using the CPU at 100%. Doing a ps on it shows that the processes are hanging at "starting children". Restarting the processes is very slow, the master process dies before the children dies. It takes ages for the children to die (>1 minute, to a maximum of 4 minutes when I ran out of patience). Restarting is the same, the processes hang at the "starting children" stage for a long time with uptime exceeding 3. Once the MailScanner process starts properly, the CPU time consumed was already more than 3:00.00 (as shown in top), I guess that's 3 hours? WOW!!! :O


The Constraints

  1. Obviously, I can't just take the email server offline and play with it.
  2. The actual problem is not obvious and really going through the source code and debugging is tough, if not impossible.
  3. MailScanner is a huge piece of software, and its not easy to find out where the process is hanging (unless Linux has something like DTrace for Solaris and assuming I know how to use it).


The Experiment

The factors which I feel are likely to affect MailScanner load and processes are listed below:
  1. MailScanner, Max Children = X
  2. MailScanner, Virus Scanning = yes|no
  3. MailScanner, Use SpamAssassin = yes|no
  4. MailScanner, spam.whitelist.rules (turn off spam checking for certain domains)
At 12-1am at night, I wasn't too awake (besides I have been coding away for the whole day), so I couldn't come out with more...

1st set of Experiments

I tried out the easiest combinations by first setting #2 #3 to "no", and then played around with #1 from 2 to 5. Nope, the only observation was that as the number of children increases, MailScanner took an (almost) exponentially longer time to start. Actually, I couldn't bother to wait and time it, I just "killall MailScanner".

2nd set of Experiments

I tried to keep the number of children, #1, constant and tested with #2 and #3 on and off alternately. Didn't help either. It seems that the problem is tied to the number of children being started.

3rd Experiment

I reinstalled MailScanner. But, it doesn't work either.

MailScanner is dependent on a lot of PERL modules. The recent server upgrade might have installed/broken something. Or it could be that the CPAN-based modules (perl -MCPAN -e shell) that I have installed previously is affected MailScanner.

One of the questions that kept bugging me is, where does CPAN installed PERL modules go, and where does RPM install PERL modules go? Which one does PERL use if both exist?


[nothing works... *sob* 2:15am... and it's all not working... so... gotta think of something fast before end-customers get online and it's DOWN, then I'll really have early morning calls with people screaming and shouting into my ear]

As a last resort, I configured the server with
  • Max Children = 2 (this still takes a couple of minutes to start)
  • Use SpamAssassin = off (but Spam List = spamhaus-ZEN is retained)
  • insert customers' domain into spam.whitelist.rules (so that outgoing emails will not be checked, and hence, this will hopefully increase the speed at which emails are relayed)
  • Restart Every = 28800 (restart every 8 hours) as the killing and respawning of children processes will cause the hang, could be lengthened to 12 hours also, since I have 1GB of RAM free
So far, so good... it's been 15 hours since...

5 hours of sleep sucks...

To really troubleshoot the problem? I installed MailScanner on a VirtualPC with CentOS 4.6 plain (no GUI, no other services except sendmail). BlueQuartz crashes when installing into a VirtualPC environment so I can't test it.

And... everything works fine in the VirtualPC!!!

ARGHH... maybe I really have to remove all the CPAN-installed modules, remove all the RPM-installed modules and stick with the ones installed by MailScanner. *sigh* Will update if this really works, what else can I do? :D