Uncle Roger has again laid down the gauntlet, establishing the January challenge for posts about resolutions:
…come up with an evaluation of last year’s goals and a new set of goals for this year.
This should not be a facile collection of cliches like “I will lose weight” or “I’m going to save some money” but carefully thought out, significant goals for your life. Don’t just list indefinite, non-specific platitudes, but specific, achievable goals. Include a plan for accomplishing each goal with concrete milestones and dates.
And:
…this challenge includes a review of your goals from last year (if any).
I’m pleased to say that I’m half finished with the challenge, as I’ve already reviewed last year’s resolutions. I’ll be posting this year’s by the end of the week(end).
These challenges have been alot of fun… I encourage you to join in this one. Be sure to link to Roger’s post, and comment over there (and here if you want). If you’ve already posted your resolutions, go link them to the challenge.
This post is for Googlebot. Go follow these links, and index them:
Thanks pal.
For the rest of you, who are probably wondering if I’ve hit my head (not that I remember), I’ll explain. While checking my referers briefly (whole other post on that topic forthcoming), I noticed what looked like a search engine hit for a topic I don’t normally post about, of an adult, or more likely a teenage male, nature. (I’m not a prude, but listing the search terms would defeat my purpose). A quick check showed that the site was a non-English search engine powered by Google.
It seems that Googlebot stopped by the day I was hit by over 2000 comment spams. Although I took the entire comment system offline to remove that crap as soon as I saw it, Googlebot must have indexed a few of the pages. Those pages linked above are the pages it indexed that day and has apparently not reindexed since. I like website traffic as much as the next blogger, but I’m really not in the market for search engine hits for animated attacks on women. I certainly don’t want to be the number #6 Google hit for that search (which I am, for the moment). The sooner Googlebot indexes the clean versions of the pages above, the better.
Also of interest, Googlebot found those pages via www.jclark.org, instead of jclark.org. They are the same, but I omit the www as it is unnecessary. Other people who link to me occaisionally use the www, which is why it is indexed both ways. Of course, it shouldn’t be indexed twice, so I’ll be adding a mod_rewrite rule to my .htaccess file to permanently redirect all www. URIs to their sub-domain-free equivalents. Just to be safe, however, I think I’ll wait until Googlebot reindexes the above links.
Inspired as I was by Michael McCraken’s I-Search plugin for Cocoa apps, I decided it was time to try another of Michael’s projects, Blapp. Blapp is a Cocoa weblog editor for Blosxom (and Blosxom-like) blogs. I’ve always edited my posts with wikieditish, a browser-based Blosxom editing solution. This has the advantage of allowing me to post from anywhere I have web access.
On the other hand, I write most of my posts on my Mac, and I can always use wikieditish when I’m not at my Mac. So far, I’m pretty impressed. Blapp allows me to preview my post as I type, even using my story template and css files from the site so I can see what the article will really look like. It also supports the use of an external app as a filter, which allows me to see my preview with full Markdown rendering.
Setup was a bit of a bear – I had to edit my story template a bit (Blapp doesn’t seem to understand interpolate_fancy-style variables, and I had to merge my css files (normally, one @import‘s the other. Figuring out the rsync setup was a bit of pain as well, but liberal command-line testing of the rsync command with the -n (make no changes) switch helped.
More good news, Blapp is open-source, so I can take a shot at fixing some of my complaints myself (when I have time, and when I finish reading Cocoa Programming for Mac OS X). If you blog with Blosxom and use OS X, I recommend giving it a try.
And yes, this is my first Blapp-edited post.
Regular readers are aware that I disabled comments on this site about a week ago, after receiving over 2000 spam comments in just a few hours. Alert readers may have also noticed that all comments have been missing since; all postings have showed 0 comments regardless of prior comments. This is because when I disabled comments, I not only wanted to prevent further postings but also to prevent the display of the 2000+ morsels of nastiness. The quickest way to do both given my Blosxom setup was to simply remove the writeback plugin.
I’ve made adjustments to my .htaccess file what will prevent the same type of posting in the future, but I expect the spammers to adapt. The delay in re-enabling comments was the need to clean out the spam. Because I use Blosxom and writeback, my comments are store in the filesystem, in a dir tree that matches the site layout, one file per blog post (i.e., all comments for a given post are in one file). With over 2000 bad comments spread across 224 different files, I wasn’t about to clean things up by hand. Instead, I wrote a perl script to help me do it.
Even though it’s used alot in Blosxom and various plugins, I’ve never had a firm grasp of perl’s File::Find module, so instead I decided to use File::Find::Rule instead. (For a nice explanation of this module, see File::Find::Rule in the 2002 Perl Advent Calendar.) The only problem is that the module (and several prerequisites) were not installed on my webserver. Having shell access, I was able to install local copies of the needed modules. Being very busy over the holidays, I only got around to this today.
The script is called scrub, and is available under the GNU General Public Licence. You can download scrub. Please note that scrub is written as a command line utility – you will need shell access on your webserver to use scrub. If there is demand (and if no one else does it first), I may develop a CGI version of scrub to run from a web server.
So how does scrub work? Why, Voodoo magic, of course. In fact, the darkest Voodoo of all… regular expressions. Supply a regex, and scrub can list all comment files containing the regex. It can also display the matching comments, and a count of files matched. Most importantly, it can remove offending comments when run with the -scrub option.
scrub is designed to work with comment files created by the writeback plugin. These files contain each comment, along with the name and url of the poster as supplied on the posting form. I’ve modified my copy of writeback to also log the IP address. Any information in the comment file can be matched by the regex… so if you are logging IPs as I am, you can quickly find (and eliminate) all comments from a given IP. scrub overrides perl’s $/ magic variable, which is the input separator. By setting $/ to "-----\n" (the comment separator in writeback files), scrub can process each comment as a single unit.
Here are a few examples of scrub usage:
-
List all files containing ‘spam.com’, display total:
scrub -regex 'spam.com' -list -count
-
Remove all comments containing ‘spam.com’, show progress via filenames:
scrub -regex 'spam.com' -list -scrub
-
Show all files containing raw html hyperlinks, and the actual comments:
scrub -regex '<a href' -list -show
Example 3 above brings to light a deficiency with scrub – the regex’s are always case sensitive. If I post a revision, this will be addressed.
If you find this at all useful, please leave me a comment… they are enabled once again.
To anyone who has viewed my site in the past 24 hours or so- If you have seen any comments on this site which you found offense, please accept my appologies. I have once again been hit by a determined comment spammer- an order of magnitude worse than anything I have seen before. Over 2000 spams have been posted. I have completely disabled the comment system.