Showing posts with label Shell Script. Show all posts
Showing posts with label Shell Script. Show all posts

Thursday, 21 March 2013

AWK One-Liners


USAGE:

Unix: awk '/pattern/ {print "$1"}' # standard Unix shells
DOS/Win: awk '/pattern/ {print "$1"}' # okay for DJGPP compiled
awk "/pattern/ {print \"$1\"}" # required for Mingw32

Most of my experience comes from version of GNU awk (gawk) compiled for
Win32. Note in particular that DJGPP compilations permit the awk script
to follow Unix quoting syntax '/like/ {"this"}'. However, the user must
know that single quotes under DOS/Windows do not protect the redirection
arrows (<, >) nor do they protect pipes (|). Both are special symbols
for the DOS/CMD command shell and their special meaning is ignored only
if they are placed within "double quotes." Likewise, DOS/Win users must
remember that the percent sign (%) is used to mark DOS/Win environment
variables, so it must be doubled (%%) to yield a single percent sign
visible to awk.

If I am sure that a script will NOT need to be quoted in Unix, DOS, or
CMD, then I normally omit the quote marks. If an example is peculiar to
GNU awk, the command 'gawk' will be used. Please notify me if you find
errors or new commands to add to this list (total length under 65
characters). I usually try to put the shortest script first.

FILE SPACING:

# double space a file
awk '1;{print ""}'
awk 'BEGIN{ORS="\n\n"};1'

# double space a file which already has blank lines in it. Output file
# should contain no more than one blank line between lines of text.
# NOTE: On Unix systems, DOS lines which have only CRLF (\r\n) are
# often treated as non-blank, and thus 'NF' alone will return TRUE.
awk 'NF{print $0 "\n"}'

# triple space a file
awk '1;{print "\n"}'

NUMBERING AND CALCULATIONS:

# precede each line by its line number FOR THAT FILE (left alignment).
# Using a tab (\t) instead of space will preserve margins.
awk '{print FNR "\t" $0}' files*

# precede each line by its line number FOR ALL FILES TOGETHER, with tab.
awk '{print NR "\t" $0}' files*

# number each line of a file (number on left, right-aligned)
# Double the percent signs if typing from the DOS command prompt.
awk '{printf("%5d : %s\n", NR,$0)}'

# number each line of file, but only print numbers if line is not blank
# Remember caveats about Unix treatment of \r (mentioned above)
awk 'NF{$0=++a " :" $0};{print}'
awk '{print (NF? ++a " :" :"") $0}'

# count lines (emulates "wc -l")
awk 'END{print NR}'

# print the sums of the fields of every line
awk '{s=0; for (i=1; i<=NF; i++) s=s+$i; print s}'

# add all fields in all lines and print the sum
awk '{for (i=1; i<=NF; i++) s=s+$i}; END{print s}'

# print every line after replacing each field with its absolute value
awk '{for (i=1; i<=NF; i++) if ($i < 0) $i = -$i; print }'
awk '{for (i=1; i<=NF; i++) $i = ($i < 0) ? -$i : $i; print }'

# print the total number of fields ("words") in all lines
awk '{ total = total + NF }; END {print total}' file

# print the total number of lines that contain "Beth"
awk '/Beth/{n++}; END {print n+0}' file

# print the largest first field and the line that contains it
# Intended for finding the longest string in field #1
awk '$1 > max {max=$1; maxline=$0}; END{ print max, maxline}'

# print the number of fields in each line, followed by the line
awk '{ print NF ":" $0 } '

# print the last field of each line
awk '{ print $NF }'

# print the last field of the last line
awk '{ field = $NF }; END{ print field }'

# print every line with more than 4 fields
awk 'NF > 4'

# print every line where the value of the last field is > 4
awk '$NF > 4'


TEXT CONVERSION AND SUBSTITUTION:

# IN UNIX ENVIRONMENT: convert DOS newlines (CR/LF) to Unix format
awk '{sub(/\r$/,"");print}' # assumes EACH line ends with Ctrl-M

# IN UNIX ENVIRONMENT: convert Unix newlines (LF) to DOS format
awk '{sub(/$/,"\r");print}

# IN DOS ENVIRONMENT: convert Unix newlines (LF) to DOS format
awk 1

# IN DOS ENVIRONMENT: convert DOS newlines (CR/LF) to Unix format
# Cannot be done with DOS versions of awk, other than gawk:
gawk -v BINMODE="w" '1' infile >outfile

# Use "tr" instead.
tr -d \r <infile >outfile # GNU tr version 1.22 or higher

# delete leading whitespace (spaces, tabs) from front of each line
# aligns all text flush left
awk '{sub(/^[ \t]+/, ""); print}'

# delete trailing whitespace (spaces, tabs) from end of each line
awk '{sub(/[ \t]+$/, "");print}'

# delete BOTH leading and trailing whitespace from each line
awk '{gsub(/^[ \t]+|[ \t]+$/,"");print}'
awk '{$1=$1;print}' # also removes extra space between fields

# insert 5 blank spaces at beginning of each line (make page offset)
awk '{sub(/^/, " ");print}'

# align all text flush right on a 79-column width
awk '{printf "%79s\n", $0}' file*

# center all text on a 79-character width
awk '{l=length();s=int((79-l)/2); printf "%"(s+l)"s\n",$0}' file*

# substitute (find and replace) "foo" with "bar" on each line
awk '{sub(/foo/,"bar");print}' # replaces only 1st instance
gawk '{$0=gensub(/foo/,"bar",4);print}' # replaces only 4th instance
awk '{gsub(/foo/,"bar");print}' # replaces ALL instances in a line

# substitute "foo" with "bar" ONLY for lines which contain "baz"
awk '/baz/{gsub(/foo/, "bar")};{print}'

# substitute "foo" with "bar" EXCEPT for lines which contain "baz"
awk '!/baz/{gsub(/foo/, "bar")};{print}'

# change "scarlet" or "ruby" or "puce" to "red"
awk '{gsub(/scarlet|ruby|puce/, "red"); print}'

# reverse order of lines (emulates "tac")
awk '{a[i++]=$0} END {for (j=i-1; j>=0;) print a[j--] }' file*

# if a line ends with a backslash, append the next line to it
# (fails if there are multiple lines ending with backslash...)
awk '/\\$/ {sub(/\\$/,""); getline t; print $0 t; next}; 1' file*

# print and sort the login names of all users
awk -F ":" '{ print $1 | "sort" }' /etc/passwd

# print the first 2 fields, in opposite order, of every line
awk '{print $2, $1}' file

# switch the first 2 fields of every line
awk '{temp = $1; $1 = $2; $2 = temp}' file

# print every line, deleting the second field of that line
awk '{ $2 = ""; print }'

# print in reverse order the fields of every line
awk '{for (i=NF; i>0; i--) printf("%s ",i);printf ("\n")}' file

# remove duplicate, consecutive lines (emulates "uniq")
awk 'a !~ $0; {a=$0}'

# remove duplicate, nonconsecutive lines
awk '! a[$0]++' # most concise script
awk '!($0 in a) {a[$0];print}' # most efficient script

# concatenate every 5 lines of input, using a comma separator
# between fields
awk 'ORS=%NR%5?",":"\n"' file



SELECTIVE PRINTING OF CERTAIN LINES:

# print first 10 lines of file (emulates behavior of "head")
awk 'NR < 11'

# print first line of file (emulates "head -1")
awk 'NR>1{exit};1'

# print the last 2 lines of a file (emulates "tail -2")
awk '{y=x "\n" $0; x=$0};END{print y}'

# print the last line of a file (emulates "tail -1")
awk 'END{print}'

# print only lines which match regular expression (emulates "grep")
awk '/regex/'

# print only lines which do NOT match regex (emulates "grep -v")
awk '!/regex/'

# print the line immediately before a regex, but not the line
# containing the regex
awk '/regex/{print x};{x=$0}'
awk '/regex/{print (x=="" ? "match on line 1" : x)};{x=$0}'

# print the line immediately after a regex, but not the line
# containing the regex
awk '/regex/{getline;print}'

# grep for AAA and BBB and CCC (in any order)
awk '/AAA/; /BBB/; /CCC/'

# grep for AAA and BBB and CCC (in that order)
awk '/AAA.*BBB.*CCC/'

# print only lines of 65 characters or longer
awk 'length > 64'

# print only lines of less than 65 characters
awk 'length < 64'

# print section of file from regular expression to end of file
awk '/regex/,0'
awk '/regex/,EOF'

# print section of file based on line numbers (lines 8-12, inclusive)
awk 'NR==8,NR==12'

# print line number 52
awk 'NR==52'
awk 'NR==52 {print;exit}' # more efficient on large files

# print section of file between two regular expressions (inclusive)
awk '/Iowa/,/Montana/' # case sensitive


SELECTIVE DELETION OF CERTAIN LINES:

# delete ALL blank lines from a file (same as "grep '.' ")
awk NF
awk '/./'


CREDITS AND THANKS:

Special thanks to Peter S. Tillier for helping me with the first release
of this FAQ file.

For additional syntax instructions, including the way to apply editing
commands from a disk file instead of the command line, consult:

"sed & awk, 2nd Edition," by Dale Dougherty and Arnold Robbins
O'Reilly, 1997
"UNIX Text Processing," by Dale Dougherty and Tim O'Reilly
Hayden Books, 1987
"Effective awk Programming, 3rd Edition." by Arnold Robbins
O'Reilly, 2001

To fully exploit the power of awk, one must understand "regular
expressions." For detailed discussion of regular expressions, see
"Mastering Regular Expressions, 2d edition" by Jeffrey Friedl
(O'Reilly, 2002).

The manual ("man") pages on Unix systems may be helpful (try "man awk",
"man nawk", "man regexp", or the section on regular expressions in "man
ed"), but man pages are notoriously difficult. They are not written to
teach awk use or regexps to first-time users, but as a reference text
for those already acquainted with these tools.

USE OF '\t' IN awk SCRIPTS: For clarity in documentation, we have used
the expression '\t' to indicate a tab character (0x09) in the scripts.
All versions of awk, even the UNIX System 7 version should recognize
the '\t' abbreviation.

#---end of file---

AWK Cheat Sheet



 ===================== Predefined Variable Summary =====================

.-------------+-----------------------------------.---------------------.
| | | Support: |
| Variable | Description '-----.-------.-------'
| | | AWK | NAWK | GAWK |
'-------------+-----------------------------------+-----+-------+-------'
| FS | Input Field Separator, a space by | + | + | + |
| | default. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| OFS | Output Field Separator, a space | + | + | + |
| | by default. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| NF | The Number of Fields in the | + | + | + |
| | current input record. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| NR | The total Number of input Records | + | + | + |
| | seen so far. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| RS | Record Separator, a newline by | + | + | + |
| | default. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| ORS | Output Record Separator, a | + | + | + |
| | newline by default. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| FILENAME | The name of the current input | | | |
| | file. If no files are specified | | | |
| | on the command line, the value of | | | |
| | FILENAME is "-". However, | + | + | + |
| | FILENAME is undefined inside the | | | |
| | BEGIN block (unless set by | | | |
| | getline). | | | |
'-------------+-----------------------------------+-----+-------+-------'
| ARGC | The number of command line | | | |
| | arguments (does not include | | | |
| | options to gawk, or the program | - | + | + |
| | source). Dynamically changing the | | | |
| | contents of ARGV control the | - | + | + |
| | files used for data. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| ARGV | Array of command line arguments. | | | |
| | The array is indexed from 0 to | - | + | + |
| | ARGC - 1. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| ARGIND | The index in ARGV of the current | - | - | + |
| | file being processed. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| BINMODE | On non-POSIX systems, specifies | | | |
| | use of "binary" mode for all file | | | |
| | I/O.Numeric values of 1, 2, or 3, | | | |
| | specify that input files, output | | | |
| | files, or all files, respectively,| | | |
| | should use binary I/O. String | | | |
| | values of "r", or "w" specify | - | - | + |
| | that input files, or output files,| | | |
| | respectively, should use binary | | | |
| | I/O. String values of "rw" or | | | |
| | "wr" specify that all files | | | |
| | should use binary I/O. Any other | | | |
| | string value is treated as "rw", | | | |
| | but generates a warning message. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| CONVFMT | The CONVFMT variable is used to | | | |
| | specify the format when | - | - | + |
| | converting a number to a string. | | | |
| | Default: "%.6g" | | | |
'-------------+-----------------------------------+-----+-------+-------'
| ENVIRON | An array containing the values | - | - | + |
| | of the current environment. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| ERRNO | If a system error occurs either | | | |
| | doing a redirection for getline, | | | |
| | during a read for getline, or | | | |
| | during a close(), then ERRNO will | - | - | + |
| | contain a string describing the | | | |
| | error. The value is subject to | | | |
| | translation in non-English locales. | | |
'-------------+-----------------------------------+-----+-------+-------'
| FIELDWIDTHS | A white-space separated list of | | | |
| | fieldwidths. When set, gawk | | | |
| | parses the input into fields of | - | - | + |
| | fixed width, instead of using the | | | |
| | value of the FS variable as the | | | |
| | field separator. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| FNR | Contains number of lines read, | - | + | + |
| | but is reset for each file read. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| IGNORECASE | Controls the case-sensitivity of | | | |
| | all regular expression and string | | | |
| | operations. If IGNORECASE has a | | | |
| | non-zero value, then string | | | |
| | comparisons and pattern matching | | | |
| | in rules, field splitting | | | |
| | with FS, record separating | | | |
| | with RS, regular expression | | | |
| | matching with ~ and !~, and the | - | - | + |
| | gensub(), gsub(), index(), | | | |
| | match(), split(), and sub() | | | |
| | built-in functions all ignore | | | |
| | case when doing regular | | | |
| | expression operations. | | | |
| | NOTE: Array subscripting is not | | | |
| | affected. However, the asort() | | | |
| | and asorti() functions are | | | |
| | affected | | | |
'-------------+-----------------------------------+-----+-------+-------'
| LINT | Provides dynamic control of the | | | |
| | --lint option from within an AWK | - | - | + |
| | program. When true, gawk prints | | | |
| | lint warnings. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| OFMT | The default output format for | - | + | + |
| | numbers. Default: "%.6g" | | | |
'-------------+-----------------------------------+-----+-------+-------'
| PROCINFO | The elements of this array | | | |
| | provide access to information | | | |
| | about the running AWK program. | | | |
| | PROCINFO["egid"]: | | | |
| | the value of the getegid(2) | | | |
| | system call. | | | |
| | PROCINFO["euid"]: | | | |
| | the value of the geteuid(2) | | | |
| | system call. | | | |
| | PROCINFO["FS"]: | | | |
| | "FS" if field splitting with FS | | | |
| | is in effect, or "FIELDWIDTHS" | | | |
| | if field splitting with | | | |
| | FIELDWIDTHS is in effect. | | | |
| | PROCINFO["gid"]: | - | - | + |
| | the value of the getgid(2) system | | | |
| | call. | | | |
| | PROCINFO["pgrpid"]: | | | |
| | the process group ID of the | | | |
| | current process. | | | |
| | PROCINFO["pid"]: | | | |
| | the process ID of the current | | | |
| | process. | | | |
| | PROCINFO["ppid"]: | | | |
| | the parent process ID of the | | | |
| | current process. | | | |
| | PROCINFO["uid"] | | | |
| | the value of the getuid(2) system | | | |
| | call. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| RT | The record terminator. Gawk sets | | | |
| | RT to the input text that matched | - | - | + |
| | the character or regular | | | |
| | expression specified by RS. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| RSTART | The index of the first character | - | + | + |
| | matched by match(); 0 if no match.| | | |
'-------------+-----------------------------------+-----+-------+-------'
| RLENGTH | The length of the string matched | - | + | + |
| | by match(); -1 if no match. | | | |
'-------------+-----------------------------------+-----+-------+-------'
| SUBSEP | The character used to separate | | | |
| | multiple subscripts in array | | | |
| | elements.Default: "\034" | - | + | + |
| | (non-printable character, | | | |
| | dec: 28, hex: 1C) | | | |
'-------------+-----------------------------------+-----+-------+-------'
| TEXTDOMAIN | The text domain of the AWK | | | |
| | program; used to find the | - | - | + |
| | localized translations for the | | | |
| | program's strings. | | | |
'-------------'-----------------------------------'-----'-------'-------'


============================ I/O Statements ===========================

.---------------------.-------------------------------------------------.
| | |
| Statement | Description |
| | |
'---------------------+-------------------------------------------------'
| close(file [, how]) | Close file, pipe or co-process. The optional |
| | how should only be used when closing one end of |
| | a two-way pipe to a co-process. It must be a |
| | string value, either "to" or "from". |
'---------------------+-------------------------------------------------'
| getline | Set $0 from next input record; set NF, NR, FNR. |
| | Returns 0 on EOF and �1 on an error. Upon an |
| | error, ERRNO contains a string describing the |
| | problem. |
'---------------------+-------------------------------------------------'
| getline <file | Set $0 from next record of file; set NF. |
'---------------------+-------------------------------------------------'
| getline var | Set var from next input record; set NR, FNR. |
'---------------------+-------------------------------------------------'
| getline var <file | Set var from next record of file. |
'---------------------+-------------------------------------------------'
| command | | Run command piping the output either into $0 or |
| getline [var] | var, as above. If using a pipe or co-process |
| | to getline, or from print or printf within a |
| | loop, you must use close() to create new |
| | instances |
'---------------------+-------------------------------------------------'
| command |& | Run command as a co-process piping the output |
| getline [var] | either into $0 or var, as above. Co-processes |
| | are a gawk extension. |
'---------------------+-------------------------------------------------'
| next | Stop processing the current input record. |
| | The next input record is read and processing |
| | starts over with the first pattern in the AWK |
| | program. If the end of the input data is |
| | reached, the END block(s), if any, are executed.|
'---------------------+-------------------------------------------------'
| nextfile | Stop processing the current input file. The |
| | next input record read comes from the next |
| | input file. FILENAME and ARGIND are updated, |
| | FNR is reset to 1, and processing starts over |
| | with the first pattern in the AWK program. If |
| | the end of the input data is reached, the END |
| | block(s), are executed. |
'---------------------+-------------------------------------------------'
| print | Prints the current record. The output record is |
| | terminated with the value of the ORS variable. |
'---------------------+-------------------------------------------------'
| print expr-list | Prints expressions. Each expression is |
| | separated by the value of the OFS variable. |
| | The output record is terminated with the value |
| | of the ORS variable. |
'---------------------+-------------------------------------------------'
| print expr-list | Prints expressions on file. Each expression is |
| >file | separated by the value of the OFS variable. The |
| | output record is terminated with the value of |
| | the ORS variable. |
'---------------------+-------------------------------------------------'
| printf fmt, | Format and print. |
| expr-list | |
'---------------------+-------------------------------------------------'
| printf fmt, | Format and print on file. |
| expr-list >file | |
'---------------------+-------------------------------------------------'
| system(cmd-line) | Execute the command cmd-line, and return the |
| | exit status. |
'---------------------+-------------------------------------------------'
| fflush([file]) | Flush any buffers associated with the open |
| | output file or pipe file. If file is missing, |
| | then stdout is flushed. If file is the null |
| | string, then all open output files and pipes |
| | have their buffers flushed. |
'---------------------+-------------------------------------------------'
| print ... >> file | Appends output to the file. |
'---------------------+-------------------------------------------------'
| print ... | command | Writes on a pipe. |
'---------------------+-------------------------------------------------'
| print ... |& | Sends data to a co-process. |
| command | |
'---------------------'-------------------------------------------------'


=========================== Numeric Functions =========================

.---------------------.-------------------------------------------------.
| | |
| Function | Description |
| | |
'---------------------+-------------------------------------------------'
| atan2(y, x) | Returns the arctangent of y/x in radians. |
'---------------------+-------------------------------------------------'
| cos(expr) | Returns the cosine of expr, which is in radians.|
'---------------------+-------------------------------------------------'
| exp(expr) | The exponential function. |
'---------------------+-------------------------------------------------'
| int(expr) | Truncates to integer. |
'---------------------+-------------------------------------------------'
| log(expr) | The natural logarithm function. |
'---------------------+-------------------------------------------------'
| rand() | Returns a random number N, between 0 and 1, |
| | such that 0 <= N < 1. |
'---------------------+-------------------------------------------------'
| sin(expr) | Returns the sine of expr, which is in radians. |
'---------------------+-------------------------------------------------'
| sqrt(expr) | The square root function. |
'---------------------+-------------------------------------------------'
| srand([expr]) | Uses expr as a new seed for the random number |
| | generator. If no expr is provided, the time of |
| | day is used. The return value is the previous |
| | seed for the random number generator. |
'---------------------'-------------------------------------------------'


====================== Bit Manipulation Functions =====================

.---------------------.-------------------------------------------------.
| | |
| Function | Description |
| | |
'---------------------+-------------------------------------------------'
| and(v1, v2) | Return the bitwise AND of the values provided |
| | by v1 and v2. |
'---------------------+-------------------------------------------------'
| compl(val) | Return the bitwise complement of val. |
'---------------------+-------------------------------------------------'
| lshift(val, count) | Return the value of val, shifted left by |
| | count bits. |
'---------------------+-------------------------------------------------'
| or(v1, v2) | Return the bitwise OR of the values provided by |
| | v1 and v2. |
'---------------------+-------------------------------------------------'
| rshift(val, count) | Return the value of val, shifted right by |
| | count bits. |
'---------------------+-------------------------------------------------'
| xor(v1, v2) | Return the bitwise XOR of the values provided |
| | by v1 and v2. |
'---------------------'-------------------------------------------------'


=========================== String Functions ==========================

.---------------------.-------------------------------------------------.
| | |
| Function | Description |
| | |
'---------------------+-------------------------------------------------'
| asort(s [, d]) | Returns the number of elements in the source |
| | array s. The contents of s are sorted using |
| | gawk's normal rules for comparing values, and |
| | the indexes of the sorted values of s are |
| | replaced with sequential integers starting with |
| | 1. If the optional destination array d is |
| | specified, then s is first duplicated into d, |
| | and then d is sorted, leaving the indexes of |
| | the source array s unchanged. |
'---------------------+-------------------------------------------------'
| asorti(s [, d]) | Returns the number of elements in the source |
| | array s. The behavior is the same as that of |
| | asort(), except that the array indices are |
| | used for sorting, not the array values. When |
| | done, the array is indexed numerically, and the |
| | values are those of the original indices. The |
| | original values are lost; thus provide a second |
| | array if you wish to preserve the original. |
'---------------------+-------------------------------------------------'
| gensub(r, s, | Search the target string t for matches of the |
| h [, t]) | regular expression r. If h is a string |
| | beginning with g or G, then replace all matches |
| | of r with s. Otherwise, h is a number |
| | indicating which match of r to replace. If t is |
| | not supplied, $0 is used instead. Within the |
| | replacement text s, the sequence \n, where n is |
| | a digit from 1 to 9, may be used to indicate |
| | just the text that matched the n'th |
| | parenthesized subexpression. The sequence \0 |
| | represents the entire matched text, as does the |
| | character &. Unlike sub() and gsub(), the |
| | modified string is returned as the result of |
| | the function, and the original target string |
| | is not changed. |
'---------------------+-------------------------------------------------'
| gsub(r, s [, t]) | For each substring matching the regular |
| | expression r in the string t, substitute the |
| | string s, and return the number of |
| | substitutions. If t is not supplied, use $0. |
| | An & in the replacement text is replaced with |
| | the text that was actually matched. Use \& to |
| | get a literal &. (This must be |
| | typed as "\\&") |
'---------------------+-------------------------------------------------'
| index(s, t) | Returns the index of the string t in the |
| | string s, or 0 if t is not present. (This |
| | implies that characterindices start at one.) |
'---------------------+-------------------------------------------------'
| length([s]) | Returns the length of the string s, or the |
| | length of $0 if s is not supplied. |
'---------------------+-------------------------------------------------'
| match(s, r [, a]) | Returns the position in s where the regular |
| | expression r occurs, or 0 if r is not present, |
| | and sets the values of RSTART and RLENGTH. |
| | Note that the argument order is the same as for |
| | the ~ operator: str ~ re. If array a is |
| | provided, a is cleared and then elements 1 |
| | through n are filled with the portions of s |
| | that match the corresponding parenthesized |
| | subexpression in r. The 0'th element of a |
| | contains the portion of s matched by the entire |
| | regular expression r. Subscripts a[n, "start"], |
| | and a[n, "length"] provide the starting index |
| | in the string and length respectively, of each |
| | matching substring. |
'---------------------+-------------------------------------------------'
| split(s, a [, r]) | Splits the string s into the array a on the |
| | regular expression r, and returns the number of |
| | fields. If r is omitted, FS is used instead. |
| | The array a is cleared first. Splitting behaves |
| | identically to field splitting. |
'---------------------+-------------------------------------------------'
| sprintf(fmt, | Prints expr-list according to fmt, and returns |
| expr-list) | the resulting string. |
'---------------------+-------------------------------------------------'
| strtonum(str) | Examines str, and returns its numeric value. |
| | If str begins with a leading 0, strtonum() |
| | assumes that str is an octal number. If str |
| | begins with a leading 0x or 0X, strtonum() |
| | assumes that str is a hexadecimal number. |
'---------------------+-------------------------------------------------'
| sub(r, s [, t]) | Just like gsub(), but only the first matching |
| | substring is replaced. |
'---------------------+-------------------------------------------------'
| substr(s, i [, n]) | Returns the at most n-character substring of s |
| | starting at i. If n is omitted, the rest of s |
| | is used. |
'---------------------+-------------------------------------------------'
| tolower(str) | Returns a copy of the string str, with all the |
| | upper-case characters in str translated to |
| | their corresponding lower-case counterparts. |
| | Non-alphabetic characters are left unchanged. |
'---------------------+-------------------------------------------------'
| toupper(str) | Returns a copy of the string str, with all the |
| | lower-case characters in str translated to |
| | their corresponding upper-case counterparts. |
| | Non-alphabetic characters are left unchanged. |
'---------------------'-------------------------------------------------'


============================ Time Functions ===========================

.---------------------.-------------------------------------------------.
| | |
| Function | Description |
| | |
'---------------------+-------------------------------------------------'
| mktime(datespec) | Turns datespec into a time stamp of the same |
| | form as returned by systime(). The datespec is |
| | a string of the form YYYY MM DD HH MM SS[ DST]. |
| | The contents of the string are six or seven |
| | numbers representing respectively the full year |
| | including century, the month from 1 to 12, the |
| | day of the month from 1 to 31, the hour of the |
| | day from 0 to 23, the minute from 0 to 59, and |
| | the second from 0 to 60, and an optional |
| | daylight saving flag. The values of these |
| | numbers need not be within the ranges |
| | specified; for example, an hour of -1 means 1 |
| | hour before midnight. The origin-zero Gregorian |
| | calendar is assumed, with year 0 preceding year |
| | 1 and year -1 preceding year 0. The time is |
| | assumed to be in the local timezone. If the |
| | daylight saving flag is positive, the time is |
| | assumed to be daylight saving time; if zero, |
| | the time is assumed to be standard time; and if |
| | negative (the default), mktime() attempts to |
| | determine whether daylight saving time is in |
| | effect for the specified time. If datespec does |
| | not contain enough elements or if the resulting |
| | time is out of range, mktime() returns -1. |
'---------------------+-------------------------------------------------'
| strftime([format | Formats timestamp according to the |
| [, timestamp]]) | specification in format. The timestamp should |
| | be of the same form as returned by systime(). |
| | If timestamp is missing, the current time of |
| | day is used.If format is missing, a default |
| | format equivalent to the output of date(1) is |
| | used. See the specification for the strftime() |
| | function in ANSI C for the format conversions |
| | that are guaranteed to be available. A |
| | public-domain version of strftime(3) and a man |
| | page for it come with gawk; if that version was |
| | used to build gawk, then all of the conversions |
| | described in that man page are available to |
| | gawk. |
'---------------------+-------------------------------------------------'
| systime() | Returns the current time of day as the number |
| | of seconds since the Epoch (1970-01-01 00:00:00 |
| | UTC on POSIX systems). |
'---------------------'-------------------------------------------------'


=============== Internationalization (I18N) Functions ================

.---------------------.-------------------------------------------------.
| | |
| Function | |
| | |
| Description | |
| | |
'---------------------+-------------------------------------------------'
| bindtextdomain(directory [, domain]) |
| |
| Specifies the directory where gawk looks for the .mo files. It |
| returns the directory where domain is ``bound.'' The default domain |
| is the value of TEXTDOMAIN. If directory is the null string (""), |
| then bindtextdomain() returns the current binding for the given domain|
'---------------------+-------------------------------------------------'
| dcgettext(string [, domain [, category]]) |
| |
| Returns the translation of string in text domain domain for locale |
| category category. The default value for domain is the current value |
| of TEXTDOMAIN. The default value for category is "LC_MESSAGES". If |
| you supply a value for category, it must be a string equal to one of |
| the known locale categories. You must also supply a text domain. Use |
| TEXTDOMAIN if you want to use the current domain. |
'---------------------+-------------------------------------------------'
| dcngettext(string1 , string2 , number [, domain [, category]]) |
| |
| Returns the plural form used for number of the translation of string1 |
| and string2 in text domain domain for locale category category. The |
| default value for domain is the current value of TEXTDOMAIN. The |
| default value for category is "LC_MESSAGES". If you supply a value |
| for category, it must be a string equal to one of the known locale |
| categories. You must also supply a text domain. Use TEXTDOMAIN if |
| you want to use the current domain. |
'---------------------'-------------------------------------------------'




=============== GNU AWK's Command Line Argument Summary ===============

.-------------------------.---------------------------------------------.
| | |
| Argument | Description |
| | |
'-------------------------+---------------------------------------------'
| -F fs | Use fs for the input field separator |
| --field-sepearator fs | (the value of the FS predefined variable). |
'-------------------------+---------------------------------------------'
| -v var=val | Assign the value val to the variable var, |
| --assign var=val | before execution of the program begins. |
| | Such variable values are available to the |
| | BEGIN block of an AWK program. |
'-------------------------+---------------------------------------------'
| -f program-file | Read the AWK program source from the file |
| --file program-file | program-file, instead of from the first |
| | command line argument. Multiple -f |
| | (or --file) options may be used. |
'-------------------------+---------------------------------------------'
| -mf NNN | Set various memory limits to the value NNN. |
| -mr NNN | The f flag sets the maximum number of |
| | fields, and the r flag sets the maximum |
| | record size. (Ignored by gawk, since gawk |
| | has no pre-defined limits) |
'-------------------------+---------------------------------------------'
| -W compat | Run in compatibility mode. In compatibility |
| -W traditional | mode, gawk behaves identically to UNIX awk; |
| --compat--traditional | none of the GNU-specific extensions are |
| | recognized. |
'-------------------------+---------------------------------------------'
| -W copyleft | Print the short version of the GNU copyright|
| -W copyright | information message on the standard output |
| --copyleft | and exit successfully. |
| --copyright | |
'-------------------------+---------------------------------------------'
| -W dump-variables[=file]| Print a sorted list of global variables, |
| --dump-variables[=file] | their types and final values to file. If no |
| | file is provided, gawk uses a file named |
| | awkvars.out in the current directory. |
'-------------------------+---------------------------------------------'
| -W help | Print a relatively short summary of the |
| -W usage | available options on the standard output. |
| --help | |
| --usage | |
'-------------------------+---------------------------------------------'
|-W lint[=value] | Provide warnings about constructs that |
|--lint[=value] | are dubious or non-portable to other AWK |
| | impl�s. With argument fatal, lint warnings |
| | become fatal errors. With an optional |
| | argument of invalid, only warnings about |
| | things that are actually invalid are |
| | issued. (This is not fully implemented yet.)|
'-------------------------+---------------------------------------------'
| -W lint-old--lint-old | Provide warnings about constructs that are |
| | not portable to the original version of |
| | Unix awk. |
'-------------------------+---------------------------------------------'
| -W gen-po--gen-po | Scan and parse the AWK program, and |
| | generate a GNU .po format file on standard |
| | output with entries for all localizable |
| | strings in the program. The program itself |
| | is not executed. |
'-------------------------+---------------------------------------------'
| -W non-decimal-data | Recognize octal and hexadecimal values in |
| --non-decimal-data | input data. |
'-------------------------+---------------------------------------------'
| -W posix--posix | This turns on compatibility mode, with the |
| | following additional restrictions: |
| | o \x escape sequences are not recognized. |
| | o Only space and tab act as field |
| | separators when FS is set to a single |
| | space, new-line does not. |
| | o You cannot continue lines after ? and :. |
| | o The synonym func for the keyword function|
| | is not recognized. |
| | o The operators ** and **= cannot be used |
| | in place of ^ and ^=.� The fflush() |
| | function is not available. |
'-------------------------+---------------------------------------------'
| -W profile[=prof_file] | Send profiling data to prof_file. |
| --profile[=prof_file] | The default is awkprof.out. When run with |
| | gawk, the profile is just a "pretty |
| | printed" version of the program. When run |
| | with pgawk, the profile contains execution |
| | counts of each statement in the program |
| | in the left margin and function call counts |
| | for each user-defined function. |
'-------------------------+---------------------------------------------'
| -W re-interval | Enable the use of interval expressions in |
| --re-interval | regular expression matching. Interval |
| | expressions were not traditionally |
| | available in the AWK language. |
'-------------------------+---------------------------------------------'
| -W source program-text | Use program-text as AWK program source |
| --source program-text | code. This option allows the easy |
| | intermixing of library functions (used via |
| | the -f and --file options) with source code |
| | entered on the command line. |
'-------------------------+---------------------------------------------'
| -W version | Print version information for this |
| --version | particular copy of gawk on the standard |
| | output. |
'-------------------------+---------------------------------------------'
| -- | Signal the end of options. This is useful |
| | to allow further arguments to the AWK |
| | program itself to start with a "-". This |
| | is mainly for consistency with the argument |
| | parsing convention used by most other POSIX |
| | programs. |
'-------------------------'---------------------------------------------'

=======================================================================

Awk One-Liners Explained, Part III: Selective Printing and Deleting of Certain Lines


4. Selective Printing of Certain Lines

45. Print the first 10 lines of a file (emulates "head -10").
awk 'NR < 11'
Awk has a special variable called "NR" that stands for "Number of Lines seen so far in the current file". After reading each line, Awk increments this variable by one. So for the first line it's 1, for the second line 2, ..., etc. As I explained in the very first one-liner, every Awk program consists of a sequence of pattern-action statements "pattern { action statements }". The "action statements" part get executed only on those lines that match "pattern" (pattern evaluates to true). In this one-liner the pattern is "NR < 11" and there are no "action statements". The default action in case of missing "action statements" is to print the line as-is (it's equivalent to "{ print $0 }"). The pattern in this one-liner is an expression that tests if the current line number is less than 11. If the line number is less than 11, Awk prints the line. As soon as the line number is 11 or more, the pattern evaluates to false and Awk skips the line.
A much better way to do the same is to quit after seeing the first 10 lines (otherwise we are looping over lines > 10 and doing nothing):
awk '1; NR == 10 { exit }'
The "NR == 10 { exit }" part guarantees that as soon as the line number 10 is reached, Awk quits. For lines smaller than 10, Awk evaluates "1" that is always a true-statement. And as we just learned, true statements without the "action statements" part are equal to "{ print $0 }" that just prints the first ten lines!
46. Print the first line of a file (emulates "head -1").
awk 'NR > 1 { exit }; 1'
This one-liner is very similar to previous one. The "NR > 1" is true only for lines greater than one, so it does not get executed on the first line. On the first line only the "1", the true statement, gets executed. It makes Awk print the line and read the next line. Now the "NR" variable is 2, and "NR > 1" is true. At this moment "{ exit }" gets executed and Awk quits. That's it. Awk printed just the first line of the file.
47. Print the last 2 lines of a file (emulates "tail -2").
awk '{ y=x "\n" $0; x=$0 }; END { print y }'
Okay, so what does this one do? First of all, notice that "{y=x "\n" $0; x=$0}" action statement group is missing the pattern. When the pattern is missing, Awk executes the statement group for all lines. For the first line, it sets variable "y" to "\nline1" (because x is not yet defined). For the second line it sets variable "y" to "line1\nline2". For the third line it sets variable "y" to "line2\nline3". As you can see, for line N it sets the variable "y" to "lineN-1\nlineN". Finally, when it reaches EOF, variable "y" contains the last two lines and they get printed via "print y" statement.
Thinking about this one-liner for a second one concludes that it is very ineffective - it reads the whole file line by line just to print out the last two lines! Unfortunately there is no seek() statement in Awk, so you can't seek to the end-2 lines in the file (that's what tail does). It's recommended to use "tail -2" to print the last 2 lines of a file.
48. Print the last line of a file (emulates "tail -1").
awk 'END { print }'
This one-liner may or may not work. It relies on an assumption that the "$0" variable that contains the entire line does not get reset after the input has been exhausted. The special "END" pattern gets executed after the input has been exhausted (or "exit" called). In this one-liner the "print" statement is supposed to print "$0" at EOF, which may or may not have been reset.
It depends on your awk program's version and implementation, if it will work. Works with GNU Awk for example, but doesn't seem to work with nawk or xpg4/bin/awk.
The most compatible way to print the last line is:
awk '{ rec=$0 } END{ print rec }'
Just like the previous one-liner, it's computationally expensive to print the last line of the file this way, and "tail -1" should be the preferred way.
49. Print only the lines that match a regular expression "/regex/" (emulates "grep").
awk '/regex/'
This one-liner uses a regular expression "/regex/" as a pattern. If the current line matches the regex, it evaluates to true, and Awk prints the line (remember that missing action statement is equal to "{ print }" that prints the whole line).
50. Print only the lines that do not match a regular expression "/regex/" (emulates "grep -v").
awk '!/regex/'
Pattern matching expressions can be negated by appending "!" in front of them. If they were to evaluate to true, appending "!" in front makes them evaluate to false, and the other way around. This one-liner inverts the regex match of the previous (#49) one-liner and prints all the lines that do not match the regular expression "/regex/".
51. Print the line immediately before a line that matches "/regex/" (but not the line that matches itself).
awk '/regex/ { print x }; { x=$0 }'
This one-liner always saves the current line in the variable "x". When it reads in the next line, the previous line is still available in the "x" variable. If that line matches "/regex/", it prints out the variable x, and as a result, the previous line gets printed.
It does not work, if the first line of the file matches "/regex/", in that case, we might want to print "match on line 1", for example:
awk '/regex/ { print (x=="" ? "match on line 1" : x) }; { x=$0 }'
This one-liner tests if variable "x" contains something. The only time that x is empty is at very first line. In that case "match on line 1" gets printed. Otherwise variable "x" gets printed (that as we found out contains the previous line). Notice that this one-liner uses a ternary operator "foo?bar:baz" that is short for "if foo, then bar, else baz".
52. Print the line immediately after a line that matches "/regex/" (but not the line that matches itself).
awk '/regex/ { getline; print }'
This one-liner calls the "getline" function on all the lines that match "/regex/". This function sets $0 to the next line (and also updates NF, NR, FNR variables). The "print" statement then prints this next line. As a result, only the line after a line matching "/regex/" gets printed.
If it is the last line that matches "/regex/", then "getline" actually returns error and does not set $0. In this case the last line gets printed itself.
53. Print lines that match any of "AAA" or "BBB", or "CCC".
awk '/AAA|BBB|CCC/'
This one-liner uses a feature of extended regular expressions that support the | or alternation meta-character. This meta-character separates "AAA" from "BBB", and from "CCC", and tries to match them separately on each line. Only the lines that contain one (or more) of them get matched and printed.
54. Print lines that contain "AAA" and "BBB", and "CCC" in this order.
awk '/AAA.*BBB.*CCC/'
This one-liner uses a regular expression "AAA.*BBB.*CCC" to print lines. This regular expression says, "match lines containing AAA followed by any text, followed by BBB, followed by any text, followed by CCC in this order!" If a line matches, it gets printed.
55. Print only the lines that are 65 characters in length or longer.
awk 'length > 64'
This one-liner uses the "length" function. This function is defined as "length([str])" - it returns the length of the string "str". If none is given, it returns the length of the string in variable $0. For historical reasons, parenthesis () at the end of "length" can be omitted. This one-liner tests if the current line is longer than 64 chars, if it is, the "length > 64" evaluates to true and line gets printed.
56. Print only the lines that are less than 64 characters in length.
awk 'length < 64'
This one-liner is almost byte-by-byte equivalent to the previous one. Here it tests if the length if line less than 64 characters. If it is, Awk prints it out. Otherwise nothing gets printed.
57. Print a section of file from regular expression to end of file.
awk '/regex/,0'
This one-liner uses a pattern match in form 'pattern1, pattern2' that is called "range pattern". The 3rd Awk Tipfrom article "10 Awk Tips, Tricks and Pitfalls" explains this match very carefully. It matches all the lines starting with a line that matches "pattern1" and continuing until a line matches "pattern2" (inclusive). In this one-liner "pattern1" is a regular expression "/regex/" and "pattern2" is just 0 (false). So this one-liner prints all lines starting from a line that matches "/regex/" continuing to end-of-file (because 0 is always false, and "pattern2" never matches).
58. Print lines 8 to 12 (inclusive).
awk 'NR==8,NR==12'
This one-liner also uses a range pattern in format "pattern1, pattern2". The "pattern1" here is "NR==8" and "pattern2" is "NR==12". The first pattern means "the current line is 8th" and the second pattern means "the current line is 12th". This one-liner prints lines between these two patterns.
59. Print line number 52.
awk 'NR==52'
This one-liner tests to see if current line is number 52. If it is, "NR==52" evaluates to true and the line gets implicitly printed out (patterns without statements print the line unmodified).
The correct way, though, is to quit after line 52:
awk 'NR==52 { print; exit }'
This one-liner forces Awk to quit after line number 52 is printed. It is the correct way to print line 52 because there is nothing else to be done, so why loop over the whole doing nothing.
60. Print section of a file between two regular expressions (inclusive).
awk '/Iowa/,/Montana/'
I explained what a range pattern such as "pattern1,pattern2" does in general in one-liner #57. In this one-liner "pattern1" is "/Iowa/" and "pattern2" is "/Montana/". Both of these patterns are regular expressions. This one-liner prints all the lines starting with a line that matches "Iowa" and ending with a line that matches "Montana" (inclusive).

5. Selective Deletion of Certain Lines

There is just one one-liner in this section.
61. Delete all blank lines from a file.
awk NF
This one-liner uses the special NF variable that contains number of fields on the line. For empty lines, NF is 0, that evaluates to false, and false statements do not get the line printed.
Another way to do the same is:
awk '/./'
This one-liner uses a regular-expression match "." that matches any character. Empty lines do not have any characters, so it does not match.

Awk One-Liners Explained, Part II: Text Conversion and Substitution


3. Text Conversion and Substitution

21. Convert Windows/DOS newlines (CRLF) to Unix newlines (LF) from Unix.
awk '{ sub(/\r$/,""); print }'
This one-liner uses the sub(regex, repl, [string]) function. This function substitutes the first instance of regular expression "regex" in string "string" with the string "repl". If "string" is omitted, variable $0 is used. Variable $0, as I explained in the first part of the article, contains the entire line.
The one-liner replaces '\r' (CR) character at the end of the line with nothing, i.e., erases CR at the end. Print statement prints out the line and appends ORS variable, which is '\n' by default. Thus, a line ending with CRLF has been converted to a line ending with LF.
22. Convert Unix newlines (LF) to Windows/DOS newlines (CRLF) from Unix.
awk '{ sub(/$/,"\r"); print }'
This one-liner also uses the sub() function. This time it replaces the zero-width anchor '$' at the end of the line with a '\r' (CR char). This substitution actually adds a CR character to the end of the line. After doing that Awk prints out the line and appends the ORS, making the line terminate with CRLF.
23. Convert Unix newlines (LF) to Windows/DOS newlines (CRLF) from Windows/DOS.
awk 1
This one-liner may work, or it may not. It depends on the implementation. If the implementation catches the Unix newlines in the file, then it will read the file line by line correctly and output the lines terminated with CRLF. If it does not understand Unix LF's in the file, then it will print the whole file and terminate it with CRLF (single windows newline at the end of the whole file).
Ps. Statement '1' (or anything that evaluates to true) in Awk is syntactic sugar for '{ print }'.
24. Convert Windows/DOS newlines (CRLF) to Unix newlines (LF) from Windows/DOS
gawk -v BINMODE="w" '1'
Theoretically this one-liner should convert CRLFs to LFs on DOS. There is a note in GNU Awk documentation that says: "Under DOS, gawk (and many other text programs) silently translates end-of-line "\r\n" to "\n" on input and "\n" to "\r\n" on output. A special "BINMODE" variable allows control over these translations and is interpreted as follows: ... If "BINMODE" is "w", then binary mode is set on write (i.e., no translations on writes)."
My tests revealed that no translation was done, so you can't rely on this BINMODE hack.
Eric suggests to better use the "tr" utility to convert CRLFs to LFs on Windows:
tr -d \r
The 'tr' program is used for translating one set of characters to another. Specifying -d option makes it delete all characters and not do any translation. In this case it's the '\r' (CR) character that gets erased from the input. Thus, CRLFs become just LFs.
25. Delete leading whitespace (spaces and tabs) from the beginning of each line (ltrim).
awk '{ sub(/^[ \t]+/, ""); print }'
This one-liner also uses sub() function. What it does is replace regular expression "^[ \t]+" with nothing "". The regular expression "^[ \t]+" means - match one or more space " " or a tab "\t" at the beginning "^" of the string.
26. Delete trailing whitespace (spaces and tabs) from the end of each line (rtrim).
awk '{ sub(/[ \t]+$/, ""); print }'
This one-liner is very similar to the previous one. It replaces regular expression "[ \t]+$" with nothing. The regular expression "[ \t]+$" means - match one or more space " " or a tab "\t" at the end "$" of the string. The "+" means "one or more".
27. Delete both leading and trailing whitespaces from each line (trim).
awk '{ gsub(/^[ \t]+|[ \t]+$/, ""); print }'
This one-liner uses a new function called "gsub". Gsub() does the same as sub(), except it performs as many substitutions as possible (that is, it's a global sub()). For example, given a variable f = "foo", sub("o", "x", f) would replace just one "o" in variable f with "x", making f be "fxo"; but gsub("o", "x", f) would replace both "o"s in "foo" resulting "fxx".
The one-liner combines both previous one-liners - it replaces leading whitespace "^[ \t]+" and trailing whitespace "[ \t]+$" with nothing, thus trimming the string.
To remove whitespace between fields you may use this one-liner:
awk '{ $1=$1; print }'
This is a pretty tricky one-liner. It seems to do nothing, right? Assign $1 to $1. But no, when you change a field, Awk rebuilds the $0 variable. It takes all the fields and concats them, separated by OFS (single space by default). All the whitespace between fields is gone.
28. Insert 5 blank spaces at beginning of each line.
awk '{ sub(/^/, "     "); print }'
This one-liner substitutes the zero-length beginning of line anchor "^" with five empty spaces. As the anchor is zero-length and matches the beginning of line, the five whitespace characters get appended to beginning of the line.
29. Align all text flush right on a 79-column width.
awk '{ printf "%79s\n", $0 }' 
This one-liner asks printf() to print the string in $0 variable and left pad it with spaces until the total length is 79 chars.
Please see the documentation of printf function for more information and examples.
30. Center all text on a 79-character width.
awk '{ l=length(); s=int((79-l)/2); printf "%"(s+l)"s\n", $0 }'
First this one-liner calculates the length() of the line and puts the result in variable "l". Length(var) function returns the string length of var. If the variable is not specified, it returns the length of the entire line (variable $0). Next it calculates how many white space characters to pad the line with and stores the result in variable "s". Finally it printf()s the line with appropriate number of whitespace chars.
For example, when printing a string "foo", it first calculates the length of "foo" which is 3. Next it calculates the column "foo" should appear which (79-3)/2 = 38. Finally it printf("%41", "foo"). Printf() function outputs 38 spaces and then "foo", making that string centered (38*2 + 3 = 79)
31. Substitute (find and replace) "foo" with "bar" on each line.
awk '{ sub(/foo/,"bar"); print }'
This one-liner is very similar to the others we have seen before. It uses the sub() function to replace "foo" with "bar". Please note that it replaces just the first match. To replace all "foo"s with "bar"s use the gsub() function:
awk '{ gsub(/foo/,"bar"); print }'
Another way is to use the gensub() function:
gawk '{ $0 = gensub(/foo/,"bar",4); print }'
This one-liner replaces only the 4th match of "foo" with "bar". It uses a never before seen gensub() function. The prototype of this function is gensub(regex, s, h[, t]). It searches the string "t" for "regex" and replaces "h"-th match with "s". If "t" is not given, $0 is assumed. Unlike sub() and gsub() it returns the modified string "t" (sub and gsub modified the string in-place).
Gensub() is a non-standard function and requires GNU Awk or Awk included in NetBSD.
In this one-liner regex = "/foo/", s = "bar", h = 4, and t = $0. It replaces the 4th instance of "foo" with "bar" and assigns the new string back to the whole line $0.
32. Substitute "foo" with "bar" only on lines that contain "baz".
awk '/baz/ { gsub(/foo/, "bar") }; { print }'
As I explained in the first one-liner in the first part of the article, every Awk program consists of a sequence of pattern-action statements "pattern { action statements }". Action statements are applied only to lines that match pattern.
In this one-liner the pattern is a regular expression /baz/. If line contains "baz", the action statement gsub(/foo/, "bar") is executed. And as we have learned, it substitutes all instances of "foo" with "bar". If you want to substitute just one, use the sub() function!
33. Substitute "foo" with "bar" only on lines that do not contain "baz".
awk '!/baz/ { gsub(/foo/, "bar") }; { print }'
This one-liner negates the pattern /baz/. It works exactly the same way as the previous one, except it operates on lines that do not contain match this pattern.
34. Change "scarlet" or "ruby" or "puce" to "red".
awk '{ gsub(/scarlet|ruby|puce/, "red"); print}'
This one-liner makes use of extended regular expression alternation operator | (pipe). The regular expression /scarlet|ruby|puce/ says: match "scarlet" or "ruby" or "puce". If the line matches, gsub() replaces all the matches with "red".
35. Reverse order of lines (emulate "tac").
awk '{ a[i++] = $0 } END { for (j=i-1; j>=0;) print a[j--] }'
This is the trickiest one-liner today. It starts by recording all the lines in the array "a". For example, if the input to this program was three lines "foo", "bar", and "baz", then the array "a" would contain the following values: a[0] = "foo", a[1] = "bar", and a[2] = "baz".
When the program has finished processing all lines, Awk executes the END { } block. The END block loops over the elements in the array "a" and prints the recorded lines. In our example with "foo", "bar", "baz" the END block does the following:
for (j = 2; j >= 0; ) print a[j--]
First it prints out j[2], then j[1] and then j[0]. The output is three separate lines "baz", "bar" and "foo". As you can see the input was reversed.
36. Join a line ending with a backslash with the next line.
awk '/\\$/ { sub(/\\$/,""); getline t; print $0 t; next }; 1'
This one-liner uses regular expression "/\\$/" to look for lines ending with a backslash. If the line ends with a backslash, the backslash gets removed by sub(/\\$/,"") function. Then the "getline t" function is executed. "Getline t" reads the next line from input and stores it in variable t. "Print $0 t" statement prints the original line (but with trailing backslash removed) and the newly read line (which was stored in variable t). Awk then continues with the next line. If the line does not end with a backslash, Awk just prints it out with "1".
Unfortunately this one liner fails to join more than 2 lines (this is left as an exercise to the reader to come up with a one-liner that joins arbitrary number of lines that end with backslash :)).
37. Print and sort the login names of all users.
awk -F ":" '{ print $1 | "sort" }' /etc/passwd
This is the first time we see the -F argument passed to Awk. This argument specifies a character, a string or a regular expression that will be used to split the line into fields ($1, $2, ...). For example, if the line is "foo-bar-baz" and -F is "-", then the line will be split into three fields: $1 = "foo", $2 = "bar" and $3 = "baz". If -F is not set to anything, the line will contain just one field $1 = "foo-bar-baz".
Specifying -F is the same as setting the FS (Field Separator) variable in the BEGIN block of Awk program:
awk -F ":"
# is the same as
awk 'BEGIN { FS=":" }'
/etc/passwd is a text file, that contains a list of the system's accounts, along with some useful information like login name, user ID, group ID, home directory, shell, etc. The entries in the file are separated by a colon ":".
Here is an example of a line from /etc/passwd file:
pkrumins:x:1000:100:Peteris Krumins:/home/pkrumins:/bin/bash
If we split this line on ":", the first field is the username (pkrumins in this example). The one-liner does just that - it splits the line on ":", then forks the "sort" program and feeds it all the usernames, one by one. After Awk has finished processing the input, sort program sorts the usernames and outputs them.
38. Print the first two fields in reverse order on each line.
awk '{ print $2, $1 }' file
This one liner is obvious. It reverses the order of fields $1 and $2. For example, if the input line is "foo bar", then after running this program the output will be "bar foo".
39. Swap first field with second on every line.
awk '{ temp = $1; $1 = $2; $2 = temp; print }'
This one-liner uses a temporary variable called "temp". It assigns the first field $1 to "temp", then it assigns the second field to the first field and finally it assigns "temp" to $2. This procedure swaps the first two fields on every line. For example, if the input is "foo bar baz", then the output will be "bar foo baz".
Ps. This one-liner was incorrect in Eric's awk1line.txt file. "Print" was missing.
40. Delete the second field on each line.
awk '{ $2 = ""; print }'
This one liner just assigns empty string to the second field. It's gone.
41. Print the fields in reverse order on every line.
awk '{ for (i=NF; i>0; i--) printf("%s ", $i); printf ("\n") }'
We saw the "NF" variable that stands for Number of Fields in the part one of this article. After processing each line, Awk sets the NF variable to number of fields found on that line.
This one-liner loops in reverse order starting from NF to 1 and outputs the fields one by one. It starts with field $NF, then $(NF-1), ..., $1. After that it prints a newline character.
42. Remove duplicate, consecutive lines (emulate "uniq")
awk 'a !~ $0; { a = $0 }'
Variables in Awk don't need to be initialized or declared before they are being used. They come into existence the first time they are used. This one-liner uses variable "a" to keep the last line seen "{ a = $0 }". Upon reading the next line, it compares if the previous line (in variable "a") is not the same as the current one "a !~ $0". If it is not the same, the expression evaluates to 1 (true), and as I explained earlier, any true expression is the same as "{ print }", so the line gets printed out. Then the program saves the current line in variable "a" again and the same process continues over and over again.
This one-liner is actually incorrect. It uses a regular expression matching operator "!~". If the previous line was something like "fooz" and the new one is "foo", then it won't get output, even though they are not duplicate lines.
Here is the correct, fixed, one-liner:
awk 'a != $0; { a = $0 }'
It compares lines line-wise and not as a regular expression.
43. Remove duplicate, nonconsecutive lines.
awk '!a[$0]++'
This one-liner is very idiomatic. It registers the lines seen in the associative-array "a" (arrays are always associative in Awk) and at the same time tests if it had seen the line before. If it had seen the line before, then a[line] > 0 and !a[line] == 0. Any expression that evaluates to false is a no-op, and any expression that evals to true is equal to "{ print }".
For example, suppose the input is:
foo
bar
foo
baz
When Awk sees the first "foo", it evaluates the expression "!a["foo"]++". "a["foo"]" is false, but "!a["foo"]" is true - Awk prints out "foo". Then it increments "a["foo"]" by one with "++" post-increment operator. Array "a" now contains one value "a["foo"] == 1".
Next Awk sees "bar", it does exactly the same what it did to "foo" and prints out "bar". Array "a" now contains two values "a["foo"] == 1" and "a["bar"] == 1".
Now Awk sees the second "foo". This time "a["foo"]" is true, "!a["foo"]" is false and Awk does not print anything! Array "a" still contains two values "a["foo"] == 2" and "a["bar"] == 1".
Finally Awk sees "baz" and prints it out because "!a["baz"]" is true. Array "a" now contains three values "a["foo"] == 2" and "a["bar"] == 1" and "a["baz"] == 1".
The output:
foo
bar
baz
Here is another one-liner to do the same. Eric in his one-liners says it's the most efficient way to do it.
awk '!($0 in a) { a[$0]; print }'
It's basically the same as previous one, except that it uses the 'in' operator. Given an array "a", an expression "foo in a" tests if variable "foo" is in "a".
Note that an empty statement "a[$0]" creates an element in the array.
44. Concatenate every 5 lines of input with a comma.
awk 'ORS=NR%5?",":"\n"'
We saw the ORS variable in part one of the article. This variable gets appended after every line that gets output. In this one-liner it gets changed on every 5th line from a comma to a newline. For lines 1, 2, 3, 4 it's a comma, for line 5 it's a newline, for lines 6, 7, 8, 9 it's a comma, for line 10 a newline, etc.