amXor

Blogging about our lives online.

5.11.2010

Copy From iPod - Mini App

0 comments

I packaged the previous post into an app. All you have to do is select your iPod folder and iPod Copy will copy it all to a folder on your desktop.

Heres the link:

iPod Copy

It's a pretty trivial Automator script, but it got me thinking about interaction. In the previous post there was about 10 lines of navigation/folder creation and one payoff line. To my friend that had never used a terminal the 10 lines were mystic voodoo, but they are really basic stuff if you've used the command line at all. And then I thought, "Whatever happened to Midnight Commander?"

Midnight Commander To The iPad

Midnight Commander enabled you to navigate visually in side-by-side views and run commands of your own or from a dropdown menu. All that painful navigation cured, but the command line was still one keystroke away. The problem was, as soon as we had visual interaction there was no going back. The average user cannot shift modes between the padded walls and unlimited undo of GUI's to the underground streetfighting of CLI's.

The announcement of the iPad is one more step into the walled garden of safe and fun computing for the masses, but it worries me that the average user knows less and less about how the technology actually functions. Even developers are becoming much more insulated from the workings of the machine. Like the Galactic Empire in Asimov's Foundation series, if we don't have knowledgable people working at every level our systems will grow unmanageable and unmaintainable.

4.22.2010

Copy From iPod

0 comments
I've had a number of requests about how to copy music from an iPod back to your computer. It's fairly straightforward if you know how to get around in Terminal, if you don't here's a walkthrough, if you do you might just want to skip ahead to the copy command, it's the only one doing anything special.
1. Open Terminal (Applications/Utilities/Terminal/)
2. Navigate to /Volumes/Your iPod/iPod_Control/Music/
 cd /Volumes/
 ls
 >>> Andrew Vanee’s iPod    BOOTCAMP    Backup    My Public Folder

 cd "Andrew Vanee’s iPod"/
 ls
 >>> Calendars Contacts  Desktop DB Desktop DF Notes    iPod_Control

 cd iPod_Control/
 ls
 >>> Artwork  Device  Music  iPodPrefs iTunes

 cd Music/
 ls
 >>> F00 F03 F06 F09 F12 F15 F18 F21 F24 F27 F30 F33 F36 F39 F42 F45 F48
 F01 F04 F07 F10 F13 F16 F19 F22 F25 F28 F31 F34 F37 F40 F43 F46 F49
 F02 F05 F08 F11 F14 F17 F20 F23 F26 F29 F32 F35 F38 F41 F44 F47

3. Make a folder on the desktop to copy to:
 mkdir ~/Desktop/ipod

4. Copy all files in folders, without Mac resource forks (-rX options)
 cp -rX ./* ~/Desktop/ipod/

5. Wait for a while. Terminal just sits there blankly, but you can go to the folder
in Finder and see the copying happening. Once you get a command prompt in 
Terminal, you're done.
Gotchas: Any folder name with spaces or weird characters needs to be surrounded by double-quotes. (eg. "Andrew Vanee's iPod"/) To do: on the fly file renaming. mdls -name kMDItemName will display the actual song name.. Not sure what to do with that.

3.08.2010

Arbor Redundancy Manager

1 comments

** UPDATE! This code will potentially create false hard-links between files that are very similar. I suddenly have many iTunes album covers that are from the wrong artist. I'm not sure if this is a result of how cavalier Apple is about messing with low level UNIX conventions, or if my code is faulty. Beware! **

We all have redundant data on our computers, especially if we have multiple computers. My new tool is aimed at making backups for this kind of thing simple.

I'll take the simplest practical example: You have two computers and one backup hard drive. You want full backup images of both of these systems. Even if you're just backing up the 'Documents' folder, there is a good chance that there is a lot of redundant files. What my little program does is checks the contents of these files, and if two are the same, it creates a hard link between them.

So, if your backup folder looks like this:

Backup
Backup > System 1 > Documents > somefile.txt
Backup > System 2 > Documents > somefile_renamed.txt

It will look exactly the same after, but there will only be one copy of the file, if the contents are identical.

I have dozens of backup CD's that contain a lot of the same information, now I don't need to sort through them and reorganize or delete duplicate files. I can leave them just as they are and any duplicate files will be linked under the hood. Here is the current state of the code, discussion to follow:

arbor.py - v. 0.02

#!/usr/bin/env python
import os, sys, hashlib

arg_error = False
if len(sys.argv) == 2:
   src = sys.argv[1]
   srcfolder = os.path.abspath(src)
   if not os.path.isdir(srcfolder):
       arg_error = True
else: arg_error = True

if arg_error:
   print "Usage: arbor [directory]"
   sys.exit()

backupfolder = os.path.join(srcfolder, ".arbor")
if not os.path.isdir(backupfolder):
   os.mkdir(backupfolder)

skipped_directories = [".Trash", ".arbor"]
skipped_files = [".DS_Store"]
size_index = {}
MAX_READ = 10485760

def addsha1file(filename, size):
   if size > MAX_READ:
       f = open(filename, 'r')
       data = f.read(MAX_READ/2)
       f.seek(size/2)
       data = data + f.read(MAX_READ/2)
       sha1 = hashlib.sha1(data).hexdigest()
   else:
       sha1 = hashlib.sha1(open(filename, 'r').read()).hexdigest()
   backupfile = os.path.join(backupfolder, sha1)
   try:
       if os.path.exists(backupfile):
           os.unlink(srcfile)
           os.link(backupfile, srcfile)
       else:
           os.link(srcfile, backupfile)
   except:
       print "Unexpected error: ", sys.exc_info()[0], sys.exc_info()[1]

fcount = 0
for root, dirs, files in os.walk(srcfolder):
   for item in skipped_directories:
       if item in dirs:
           dirs.remove(item)
   for name in files:
   fcount += 1
   if fcount % 500 == 0:
       print fcount, " files scanned"
       if name in skipped_files:
           # print name
           continue

       srcfile = os.path.join(root, name)
       size = os.stat(srcfile)[6]
       if size_index.has_key(size):
           if size_index[size] == '':
               addsha1file(srcfile, size)
           else:
               addsha1file(size_index[size], size)
               addsha1file(srcfile, size)
               size_index[size] = ''
       else:
           size_index[size] = srcfile

Here are the added features of this version:

  • Folder to backup is passed as command-line argument
  • Backup files are placed at the top level of that folder in the .arbor directory. These are just hard links so they don't really add any to the size of the directory.
  • Ability to skip named folders or files
  • Only calculates a checksum if two files are the same size.
  • Only calculates a partial checksum if a file is over 10mb. Checks 5mb from start and 5mb from middle.
  • Prints a running tally of files checked (per 500 files)
  • Doesn't choke on errors: some files don't like to be stat'ed or unlinked, permissions issues.

I'm not sure about the partial checksum option, but it was really bogging down on larger inputs. It's not really practical to do a SHA1 checksum on a bunch of large files, and i think it's safe to say that two very large files can be assumed to be the same if the first 5mb, the middle 5mb and the overall size are exactly the same. Perhaps I will add an option later for strict checking, if someone is highly concerned about data integrity. But the practical limitations are there, I'm 30,000 files into a scan of my ~200GB backup folder and I certainly wouldn't have gotten that far without the file size limiting.

Update: The scan was almost done when I wrote this. Here is the tail end of the log:

31000  files scanned
31500  files scanned
32000  files scanned
32500  files scanned

real 43m40.531s
user 7m25.593s
sys 3m35.540s

So 233 GB over 32500 items took about 45 minutes to check and it looks like I've saved about 4 GB. Upon further inspection, it seems that most media files save their metadata in the file contents, so the checksum is different. Hmmm....

3.06.2010

File Management Tool - Part 2

0 comments

I have worked with my program a bit more and there are some interesting aspects of this kind of backup.

At the end of the backup all duplicate files point to the same inode, so the SHA1 version can be deleted.

eg. 

inode   name
1299    workingdir/folder1/file1.txt
1299    workingdir/folder5/file1.txt
1299    workingdir/folderx/file1_renamed.txt
1299    backupdir/54817fa363dc294bc03e4a70f51f5411f4a0e9a9

All these files now point at the same inode and so the backup directory can be erased and no file has executive control over this inode. All three files would have to be deleted to finally get rid of inode 1299. Generally it seems that programs save files with new inodes (Text Edit ...), so editing any of the versions breaks the links. It seems that UNIXy programs respect the inode better, vim saves with the same inode and so editing any version edits every version.

Removing the "backup" directory also helps Spotlight resolve the names and filetypes. Deleting that folder and running `mdimport ./workingdir` complained mightily but more or less re-indexed the folder. Here is a quick slice of the errors it produced, I'm not going to try to make sense of them, but think they're interesting; maybe Spotlight encounters these kinds of problems always and just keeps silent about them.

$mdimport ./workingdir
...
font `F88' not found in document.
font `F82' not found in document.
font `F88' not found in document.
font `F82' not found in document.
font `F88' not found in document.
font `F82' not found in document.
font `F88' not found in document.
encountered unexpected symbol `c'.
encountered unexpected symbol `c'.
encountered unexpected symbol `c'.
encountered unexpected symbol `c'.
encountered unexpected symbol `c'.
encountered unexpected symbol `c'.
encountered unexpected symbol `c'.
choked on input: `144.255.258'.
choked on input: `630.3.9'.
choked on input: `681.906458.747'.
choked on input: `680.335458.626'.
choked on input: `682.932458.507'.
choked on input: `530.3354382.624'.
font `Fw' not found in document.
font `Fw8' not found in document.
encountered unexpected symbol `w6.8'.
encountered unexpected symbol `w0.5'.
font `Fw8' not found in document.
encountered unexpected symbol `w0.5'.
encountered unexpected symbol `w6.8'.
choked on input: `397.67.'.
choked on input: `370.5g'.
choked on input: `370.5g'.
choked on input: `D42.32 m
314.94 742.32 l
S
306.06 751.2 m
306.06 7...'.
choked on input: `67.l'.
choked on input: `67.l4'
failed to find start of cross-reference table.
missing or invalid cross-reference trailer.

To reiterate, this is a funny trick that my program is doing. It builds a list of SHA1 named files from the source directory and then you just delete the index it just made and you're left with all the duplicates hard linked. I think that's pretty cool.

Metadata

One stated aim of this backup tool was to preserve metadata. So far this tool preserves the time stamps and metadata of whatever file it indexes first and the filename of every file it indexes. I'm not sure how to implement any more than this in a transparent way. As far as I can tell from the documentation, you can't have a single inode with multiple access and modification times. And building an external database of that kind of information would not get used.

File Management Tool

0 comments

I ended up sketching out the details of how my file managing tool will work, it's kind of like a virtual librarian that removes redundant files without deleting the file hierarchy. My method is a bit of a mashup of how other tools work, so I'll give credit where credit is due.

This is how it works so far:

  • All files in a tree have their SHA1 hash-value computed (Git)
  • A hard link is created in the backup folder with the SHA1 name (Time Machine) ...
  • unless: the file exists already, then it is hard-linked to the existing SHA1 (...)

There is no copying or moving of files, simply linking and unlinking, so 99.9999% of the time is spent computing the hash values of the files. Here's the python version of this:

import os
import hashlib

backupfolder = os.path.abspath('./backup')
srcfolder = os.path.abspath('./working')
srcfile = ''
backupfile = ''

for root, dirs, files in os.walk(srcfolder):
    for name in files:
        if name == '.DS_Store':
            continue
        srcfile = os.path.join(root, name)
        sha1 = hashlib.sha1(open(srcfile, 'r').read()).hexdigest()
        backupfile = os.path.join(backupfolder, sha1)
        if os.path.exists(backupfile):
            os.unlink(srcfile)
            os.link(backupfile, srcfile)
        else:
            os.link(srcfile, backupfile)
        # print backupfile

This folder contains about 5 GB of info and I thought that the SHA1 calculations might take a couple weeks, but as it turns out, it only takes a couple minutes. What you end up with is a backup folder that contains every unique file within this tree named by it's sha1 tag, and the source folder looks exactly as when you started, but every file is a hard link.

So, what are the benefits?

Filenames are not important

Because the SHA1 only calculates the contents of a file, filenames are not important. This is important in two ways, if a file has been renamed in one tree, yet remains physically the same, you only have one copy and the unique names are preserved. And more importantly, if you have two files in separate trees that are named the same, (ie. 'Picture 1.png'), you keep the naming, yet have different files.

If you have some trees of highly redundant data, this is the archive method for you. My test case was a folder of 15 direct copies of backup CD's that I have made over the years and I have saved about 600M across 5GB. And the original file hierarchies look exactly the same as they did before running the backup.

What is wrong with it?

As it stands, it messes with Spotlight and Finder's heads a little bit. Finder isn't computing correct size values for the two folders. du prints the same usage whether I include both folders or one at a time, which is pretty clever: total:5.1GB, working:5.1GB, backup:5.1GB. Finder on the other hand prints Total: 5.1GB, working: 5.1GB, backup: 4.22GB.

Spotlight

Some very wierd stuff happens with Spotlight.

A Spotlight search in the working directory will show mostly files from the backup directory, which isn't convenient. The files in the backup dir have no file extension so they're essentially unopenable by Finder. Here's what i found using the command-line mdfind:

mdfind -onlyin ./working "current"
/Users/.../backup/5f5b587eb07ee61f15ab0a032ca564a17ff461e9
/Users/.../backup/0f3f769000f164b2e30bb7b3f09482e8cc244135
and so on ...

mdfind -onlyin ./backup "current"
nothing found

For some reason, when searching the working directory it finds the information, yet always resolves the name of the file to a directory it's not supposed to be searching. And if you search the backup directory, it doesn't even bother reading the files, because it assumes from the name that they are unreadable by it.

I'm starting to wish that Steve Jobs hadn't caved and given in to the file extension system.

Time Machine

Okay, it's useful but how is it similar to Time Machine? Time Machine creates a full copy of the tree when it first backs up the system. From then on it creates the full hierarchy of directories but all the files that haven't changed are hard links to the original backup. Each unique file is a new inode created in time, whereas in my system each unique file is a new inode created in space. All duplicates in time are flattened by Time Machine and all duplicates in space are flattened by my system.

Note: To copy folders from the command line and preserve as much as possible for metadata use `cp -Rp`.

3.05.2010

File Backup And Synchronization

0 comments
In my previous post I had mentioned that I was looking for a backup/file synchronization tool.
I don't think Git is it and neither is dropbox. Both these are useful in that they are format transparent, which most database software is not. But what they are lacking is a way to deal with a large variety of file and folder hierarchies and seamlessly compress without losing transparency and semantic meaning.
So here is my list of requirements from a backup tool:
  1. Preserves any time-stamp information, even conflicting
  2. Distributed (decentralized)
  3. Minimizes redundant data
  4. Preserves hierarchies for semantic meaning
  5. Hides hierarchy clutter
  6. Preserves every bit of metadata, even if it's not explicit
  7. Accessible and platform neutral
  8. Makes data integrity paramount
It may seem like I have requirements that conflict with each other, but I will try to explain what I mean. I have loaded four of my backup CD's onto my laptop. I know there are duplicate files and I know there are time-stamps that disagree with one another.
I want to be able to view these files in a number of ways:
  • In their original on-disk hierarchy.
  • By file type, date, tags or physical description.
And I want to be able to synchronize all or part of these folders between machines, in addition to making zip/tar archive of them to a backup machine.
Any suggestions, or shall I start coding?

Twitter

Labels

Followers

andyvanee.com

Files